The AI productivity gap is the spread between what most people get from AI and what the top performers get. On August 21, 2026, McKinsey put a number on it: roughly 80 percent of the software engineers it surveyed reported an average AI productivity acceleration of approximately 3 percent, while the top 20 percent reported an average of 55 percent. Same survey, same period. Most organizations report neither figure. They report the blended average, and the blended average of that distribution describes almost nobody.
How big is the AI productivity gap?
The research is drawn from a May 2026 survey of 334 product and engineering respondents across five regions, including a subset of 173 at director level or above. The developer analysis behind the 3 percent and 55 percent figures carries a sample of 198.
Three findings sit next to each other and are worth reading together. Only 25 percent of director-and-above respondents reported meaningful acceleration, which McKinsey defines as more than a quarter of their teams achieving twofold or greater productivity gains. Thirty percent reported that team productivity had fallen. And the engineer-level spread runs from about 3 percent for the bulk of the population to about 55 percent for the top quintile.
The spread is wider than any single number
McKinsey's own framing is competitive: some organizations are pulling ahead while others see little impact or negative returns. What it does not report, and what almost no company can report about itself, is how much of that spread sits between organizations and how much sits inside one. The top quintile's reported gain is roughly eighteen times the gain the other four fifths report. Until an organization can locate itself on that distribution, it does not know whether it is holding a leading average or a wide internal spread.
Read the sample before you read the spread
This is self-reported survey data, sampled in the hundreds rather than the thousands, fielded three months before publication. Reasonable evidence for direction, weak evidence for precision. The claim worth carrying forward is not the exact 3 to 55 ratio. It is that the distribution is wide enough that no single company-level number can represent it.
Why does a company-level AI productivity number hide this?
Because an average is a compression, and compression is lossy in a predictable direction. It pulls the tails toward the middle and returns a figure that is defensible in a board deck and unusable in an operating review.
The 30 percent of director-level respondents who reported falling team productivity are the sharper problem. In a mean, a negative outcome does not appear as a negative outcome. It appears as a slightly smaller positive one. An organization can hold a credible company-wide gain of 13 percent while a third of its team leaders watch their teams move backwards, and nothing in the reported number will say so.
Our earlier analysis of AI cost reduction made a related argument one level up: a cost cannot be removed from an operating model that cannot see where the cost moved. Distribution asks the prior question. Whose gain, and did anyone else get it.
Does the gap show up outside engineering?
It does, and the second dataset is useful because it comes from a different function and a different vendor. SolarWinds surveyed more than 800 technology professionals for its 2026 State of ITSM report, covered by CIO Dive on August 20, 2026. Respondents saved roughly 3.3 hours a week detecting and flagging issues, 3 hours on end-user requests and 2.9 hours triaging tickets and knowledge base articles.
The same research reported that 71 percent had seen overall workload stay flat or rise, 52 percent had seen it rise outright, and 84 percent said AI had met or exceeded ROI expectations.
Work moved, it did not evaporate
The mechanism is in the same release: 48 percent said they now spend time managing and maintaining AI tools and integrations, 47 percent spend more time reviewing and validating AI-generated output, and 37 percent spend time training and fine-tuning models. This is vendor-published research and the fielding window is not disclosed, so the magnitude deserves caution. Hours are returned at the task level and reabsorbed at the role level, and the net effect varies by person.
Verification is a large part of why the median moves so little. Stack Overflow's 2025 Developer Survey found only 3.1 percent of respondents highly trust the accuracy of AI output, with 46 percent distrusting it against 33 percent trusting it. People who do not trust an output check it. Checking is work, and it is invisible to every dashboard that counts prompts.
What separates the top quintile?
Not the tool. McKinsey's four themes across the organizations pulling ahead are process redesign, role redesign, verification systems that keep pace with faster work, and sustained investment in change management. The measurement finding matters most here: 86 percent of top-accelerating organizations track outcome metrics such as quality, productivity and speed, while low-accelerating organizations disproportionately rely on tool adoption as their primary indicator.
The pattern is consistent. Organizations that embedded AI and redesigned roles and ways of working, which McKinsey calls transformers, became top accelerators 40 percent of the time. Organizations that embedded the tools without changing roles reached that bar 27 percent of the time. Organizations that adopted tools without embedding them in the workflow reached it 17 percent of the time.
One case in the same research makes the point concrete. At a global bank, early copilot adoption delivered 10 to 15 percent productivity improvement and little new value. After a structured program that seconded experienced AI engineers into delivery squads and wrote throughput targets into team objectives, in-scope teams reached efficiency gains of 40 to 80 percent. Same category of tool. Different operating system around it.
How adoption metrics and outcome metrics compare
| Question | Adoption metrics answer | Outcome metrics answer |
|---|---|---|
| Licenses, daily active users, prompt volume | Is the tool being touched | Nothing |
| Share of code generated by AI | How much output the model produced | Nothing about whether it shipped |
| Cycle time, rework rate, defect resolution | Nothing | Whether the work got faster and safer |
| Team-level variance against a control group | Nothing | Which teams moved and under what conditions |
| Who is in the top quintile and why | Nothing | The finding that transfers to other teams |
McKinsey's survey also found average time savings of 11.8 percent against average rework reduction of 6.2 percent, and concluded that speed is improving faster than quality. An adoption metric cannot see that divergence at all.
What measuring the AI productivity gap actually requires
The reason most organizations report an average is not laziness. It is that the average is the only number their systems can produce. Performance systems collect manager opinion on a quarterly cycle. Finance systems collect invoices. Neither observes the work.
In a 2024 round of multi-organization workforce conversations, we heard this repeatedly. At one global pharmaceutical manufacturer of roughly 100,000 employees, a modern HR stack still sat underneath goal entry done by hand, with no automation in tracking. When goals are typed manually and reviewed twice a year, nothing in the stack can distinguish a team that gained 55 percent from a team that lost ground.
Four requirements, none of them exciting
- A baseline measured before the capability arrived, not reconstructed afterward.
- A comparison group of similar teams that did not adopt it.
- Controls for tenure, role mix, work type and tool stack.
- A disclosed confidence level, including the cases where confidence is low.
This is what Levos means by controlled cohort analysis with confidence scoring. It does not claim full attribution and we say so in public. What it produces is not a headline percentage but a distribution: which teams moved, by how much, under what conditions, and whether those conditions can be copied.
That is also why the AI Impact signal family treats adoption as a behavioral signal drawn from the tools where work happens rather than as a license count, and why skills and high-potential identification run off demonstrated work rather than self-report. If the top quintile's gain is roughly eighteen times what everyone else reports, the highest-value question in the organization is who is in that quintile and what they do differently. An adoption dashboard cannot answer that.
Frequently asked questions
What is the AI productivity gap?
The spread between what most employees get from AI and what the top performers get. McKinsey's May 2026 survey of 334 product and engineering respondents found roughly 80 percent of engineers reporting about 3 percent acceleration and the top 20 percent reporting 55 percent. McKinsey does not split that spread into between-company and within-company variance, and almost no organization holds the data to do it for itself.
Why do company-level AI productivity numbers look flat?
Because a mean across a skewed distribution is correct and useless. If four fifths move 3 percent and one fifth moves 55 percent, the average lands around 13 percent and describes nobody. It also hides losses: 30 percent of McKinsey's director-level respondents reported team productivity had fallen, which a blended figure records only as a slightly smaller gain.
Do AI time savings reduce workload?
Often not. SolarWinds surveyed more than 800 technology professionals for its 2026 State of ITSM report, covered by CIO Dive on August 20, 2026. Respondents saved roughly 3 hours a week per task category, while 71 percent said overall workload stayed flat or rose and 84 percent still said AI met or exceeded ROI expectations. Both are true. The work moved rather than disappeared, with 47 percent reporting more time reviewing AI output. Vendor research, fielding window not disclosed.
What do the highest-performing organizations measure differently?
Outcomes rather than usage. McKinsey found 86 percent of top accelerators track outcome metrics such as quality, productivity and speed, while low accelerators lean on tool adoption. Organizations that redesigned roles alongside the technology became top accelerators 40 percent of the time, against 27 percent for tool embedding alone and 17 percent for adoption alone.
How do you measure the AI productivity gap credibly?
Four things most stacks cannot supply: a baseline from before the capability arrived, a comparison group of similar teams that did not adopt it, controls for tenure, role mix, work type and tool stack, and a disclosed confidence level. Levos calls this controlled cohort analysis with confidence scoring. It does not claim full attribution, and it returns a distribution rather than a headline percentage.
The next step
If your AI productivity number is a single percentage, you are reporting the mean of a distribution you have never seen. The first useful move is not a better estimate. It is finding out how wide the spread is, and who sits at each end.
Request a Demo to see how Levos measures the distribution rather than the average, or read how AI adoption and outcome signals are derived from the tools where work happens.
Levos is accepting design partner applications from US organizations of 150 to 2,500 employees with an active AI rollout.