Skip to Content

Why AI ROI Studies Contradict Each Other, and Which One Your CFO Should Believe

July 20, 2026 Tomer Mann 7 min read
Why AI ROI Studies Contradict Each Other, and Which One Your CFO Should Believe

If you have tried to answer the question "did our AI investment work," you have probably noticed that the research does not agree with itself.

Deloitte's enterprise research reports that a large majority of organizations say their most advanced generative AI initiative is meeting or exceeding ROI expectations. MIT's NANDA initiative, examining more than 300 publicly disclosed AI initiatives, reports that 95 percent of generative AI pilots showed no measurable impact on profit and loss.

Both studies fielded roughly the same kind of organization in roughly the same window. Both cannot be describing the same reality.

The reason has nothing to do with sampling error, and it sits underneath almost every AI ROI number currently in circulation.

What is definitional drift in AI ROI measurement?

Definitional drift is what happens when the ROI number moves with the definition of success rather than with the work itself.

Deloitte asked respondents to evaluate their initiative against the respondent's own expectations. MIT NANDA evaluated initiatives against an external fixed standard, measurable profit-and-loss impact.

That is the entire difference between the two headlines.

When the standard is loose and self-set, AI looks like a broad success. When the standard is fixed and applied from outside, the same population looks stuck. It is not a rounding error. It is the difference between a success story and a failure story, and the gap is an artifact of who was allowed to define success.

How the two framings compare

  Self-set standard External fixed standard
Who defines success The respondent The researcher
Typical question Did this meet your expectations Did this move P&L
Representative finding Majority meeting or exceeding ROI expectations (Deloitte) 95 percent of pilots showed no measurable P&L impact (MIT NANDA)
What it is useful for Sentiment and internal alignment Board-defensible attribution
What it cannot survive An audit Nothing, but it is harder to collect

Note on the 95 percent figure: it has been contested on sample-composition grounds. It belongs inside a triangulated set, never as a standalone statistic. The argument here holds without it.

What do the independent reads actually show?

Strip out the self-graded numbers and the picture tightens considerably.

IBM's CEO study found that about 25 percent of AI initiatives delivered the ROI expected of them, and that about 29 percent of executives say they can confidently measure AI ROI against 79 percent who already perceive gains.

Sit with that second pair. Roughly eight in ten leaders believe they are getting value. Roughly three in ten can demonstrate it. That 50-point spread is the actual state of enterprise AI measurement in 2026.

PwC's 29th Annual Global CEO Survey, covering 4,454 CEOs across 95 countries, found 56 percent report no significant financial benefit from AI to date.

McKinsey's State of AI 2025, with 1,993 respondents fielded June to July 2025, found 88 percent of organizations now use AI in at least one business function, up from 78 percent the prior year, while only about 6 percent qualify as high performers attributing 5 percent or more of EBIT to AI.

Adoption is wide. Attribution is narrow. Most measurement programs quietly fail in the space between those two facts.

Why can employees not tell you how much time AI saves them?

The most useful study in this area is also the smallest.

METR ran a randomized controlled trial with 16 experienced open-source developers across 246 real issues. Developers using AI tools were 19 percent slower on familiar work. The same developers self-reported being 20 percent faster. A 39-point gap between what happened and what people believed happened.

The sample is small and the setting is specific. Neither of those facts rescues the survey-based approach.

If skilled practitioners cannot estimate their own throughput change on work they do every day, a quarterly engagement survey asking employees how much time AI saves them is not producing a measurement. It is producing a sentiment reading. Sentiment is not what a CFO is asked to defend.

Does this only affect engineering teams?

No. Developer studies get the attention because developers generate the cleanest telemetry. Pull requests, commits, and defect rates are easy to count. The measurement problem is worse everywhere else, not better.

A claims operations example

Consider a claims operations group inside a 6,000-employee insurance carrier that deploys an AI assistant for first-line triage. Claims per handler rise 11 percent.

Then the harder questions arrive.

The three questions that decide whether the number is real

Did reopen rate move with it? A throughput gain paid for with rework is not a gain.

Did the queue mix shift toward easier claims in the same window? Composition change is the most common false positive in operational measurement.

Did the three comparable claims teams in the same organization move the same way for reasons unrelated to AI? Without that comparison, there is no way to separate the AI effect from everything else that changed.

Ask the team how much time the assistant saves them, and you will get a number. It will not be the number. The same structure applies to FP&A close cycles and support deflection.

What does a defensible AI ROI measurement require?

If self-grading and self-report are both out, what remains is not complicated to describe, though it is demanding to run.

  1. Fix the definition before you measure. The standard has to be set externally and in advance. A team's pre-stated quarterly objectives are a defensible bar. A retrospective judgment that things feel better is not.
  2. Measure behavior across the whole tool stack. Work does not happen in one place. A document drafted in one assistant is reviewed in a second system, approved in a third, and only shows up as an outcome in a fourth. Any single-tool view misses the cascade by construction. This is the gap the Levos platform is built to close.
  3. Compare against a matched control inside the same organization. A single team's before-and-after is not a measurement, because hiring shifted, the market moved, and a reorganization probably happened in the same window. Controlled cohort analysis matches the measured team against comparable teams on tenure distribution, role mix, project class, and tool stack, then reports the residual with a confidence score attached. It removes observable confounders. It does not remove unobservable ones, and any honest method says so plainly.
  4. Publish a reliability score and show the arithmetic. A reviewer should be able to recompute the result from disclosed inputs without contacting the vendor. A number you cannot reproduce is a dashboard. A number you can reproduce is a measurement. The full decomposition is documented in the Levos measurement methodology.
  5. Check the result against macroeconomic reality. Daron Acemoglu's analysis caps AI's total factor productivity contribution at not more than 0.66 percent over 10 years, substantially below what vendors routinely claim from individual deployments. Any team-level result implying an economy-wide effect several times larger than the most rigorous published macro estimate owes the reader an explanation.

Two things are true at once

AI is producing real gains in specific, well-instrumented settings. And the majority of organizations cannot currently prove it either way.

Both of those statements are supported by the evidence above, and the tension between them is the whole problem. The organizations that resolve it will be the ones that fixed the definition before they built the dashboard.

We published our entire measurement methodology for exactly this reason. Twelve pages covering the six signal families, the cohort matching specification, the confidence score arithmetic with two worked examples, sector benchmarks built as residuals above BLS labor productivity baselines, and a full page on what the method cannot do. It is free and it is written to be argued with.

Request a Demo to see the methodology applied to your own cohorts, or read the measurement methodology first and decide whether the method holds up before you talk to anyone.

Frequently asked questions

What is definitional drift in AI ROI measurement?

Definitional drift is when a reported ROI figure changes based on who defined success rather than on what the work produced. Studies that let respondents grade themselves against their own expectations report far higher AI success rates than studies applying a fixed external standard such as measurable profit-and-loss impact, even when both survey similar organizations in the same period.

Why do Deloitte and MIT report opposite findings on AI ROI?

Deloitte asks organizations to evaluate their AI initiative against their own expectations, which produces a majority reporting success. MIT NANDA applied an external standard of measurable P&L impact across more than 300 disclosed initiatives and found 95 percent showed none. The divergence comes from the standard applied, not from the underlying organizations.

Can employees accurately report how much time AI saves them?

Generally no. METR's randomized controlled trial found experienced developers were 19 percent slower with AI on familiar work while reporting themselves 20 percent faster, a 39-point gap between measured and perceived productivity. Self-reported time savings should be treated as sentiment data rather than measurement.

What is controlled cohort analysis?

It compares a team using AI against matched comparison teams inside the same organization, matched on tenure distribution, role mix, project class, and tool stack. The residual difference, reported with a confidence score, is the defensible estimate of AI-attributable change. It is observational, so unmeasured confounders remain unmeasured, and Levos discloses that limit in every published output.

How should a CFO evaluate an AI measurement vendor?

Ask whether the methodology is published, whether a published number can be recomputed from disclosed inputs without contacting the vendor, whether the vendor measures its own product or the customer's full tool stack, and whether the vendor discloses what its method cannot do. A vendor that will not publish its method is asking for trust it has not earned.

Deloitte. "State of AI in the Enterprise 2026." n = 3,235 leaders across 24 countries. https://www2.deloitte.com/us/en/pages/about-deloitte/articles/press-releases.html

MIT NANDA. "The GenAI Divide: State of AI in Business 2025." July 2025. More than 300 publicly disclosed AI initiatives, 52 interviews, 153 senior leader surveys. https://nanda.media.mit.edu/ai_report_2025.pdf

IBM Institute for Business Value. "IBM CEO Study 2025 to 2026." n = 2,000 CEOs across 33 countries and 24 industries, fielded February to April 2025. https://newsroom.ibm.com/2025-05-06-ibm-study-ceos-double-down-on-ai-while-navigating-enterprise-hurdles

PwC. "29th Annual Global CEO Survey." January 19, 2026. n = 4,454 CEOs across 95 countries. https://www.pwc.com/gx/en/news-room/press-releases/2026/pwc-2026-global-ceo-survey.html

McKinsey. "The State of AI 2025." n = 1,993, fielded June to July 2025.

METR. "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." July 10, 2025. n = 16 developers, 246 issues. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

Acemoglu, Daron. "The Simple Macroeconomics of AI." NBER Working Paper 32487. Economic Policy 40(121), 2025. https://www.nber.org/papers/w32487

U.S. Bureau of Labor Statistics. "Productivity and Costs by Industry." https://www.bls.gov/news.release/prod2.nr0.htm

Share this article

Help others discover workforce intelligence insights

Levos Editorial

Levos Editorial publishes operator-grade research on workforce intelligence, AI deployment measurement, and human capital optimization. Reach the team at marketing@levos.ai