Gartner's September 1 2026 release defines an AI high performer by four operating practices, not by models, budgets or talent. It then reports that the companies meeting that definition saw positive returns on 81% of their AI initiatives. The definition is the interesting part. Three of the four practices are common. The fourth is rare, and it is the one that makes the other three worth doing.
Two things are true at once here. Measurement discipline is not a glamorous answer, and the evidence for it is definitional rather than experimental. Gartner grouped companies by practice and then measured what they reported. That is worth reading closely and it is not worth overstating.
What Gartner actually found
Gartner surveyed 1,303 respondents at organizations with at least $50M in enterprisewide revenue in fiscal 2025. The survey was fielded January to April 2026 and published on September 1. Those are different dates and the fielding window is the one that matters.
The headline is that only 22% of organizations have successfully scaled AI across multiple business units or adopted an AI-first approach. Spending is rising anyway. In Gartner's words, "Eighty-five percent of functional leaders plan to increase spending in 2026 after dedicating an average of 12% of their functional budgets to AI in 2025."
Gartner also found roughly 11% of organizations are entirely unaware of what their function spent on AI in 2025. Note the scope: that is what a function spent, not what the organization spent. Tina Nunno, Distinguished Vice President and Gartner Fellow, connects the two: "This lack of financial visibility heightens risk as spending accelerates. Without disciplined measurement tied directly to business outcomes, organizations risk wasted resources and unmet expectations."
The useful part is the definition. Gartner describes high performers as "companies that constantly track ROI of AI initiatives, treat AI as a portfolio of value and regularly assess project performance and reallocate or discontinue underperforming initiatives." Those companies "reported positive returns in 81% of their AI initiatives, while low performers reported they did not know the rate of return for 29% of AI initiatives."
The sentence people are about to misread
That construction invites an inference Gartner does not make, so it is worth stopping on.
Gartner defines high performers. It does not define low performers as the companies that fail to track. It reports what one group does and what the other said, leaving the space between them empty. Read carelessly, the passage becomes "not tracking ROI means you lose sight of the return on nearly 30% of your AI initiatives." That is a claim about cause and it is not in the source.
We quote it rather than paraphrase it for that reason. The finding is Gartner's and the reader can draw their own line. A measurement company that overstates a measurement finding has a problem no amount of good writing fixes.
The four practices, and what each one requires
Unpacked, the definition is an operating model rather than a sentiment.
| Practice | What it means | What it requires to be real |
|---|---|---|
| Constantly track ROI | Continuous, not an annual retrospective | Spend visibility at the level being reviewed |
| Treat AI as a portfolio of value | Holdings with different risk and return, not a project list | A shared unit of comparison across initiatives |
| Regularly assess performance | On a schedule, not when someone asks | A defined outcome measure per initiative |
| Reallocate or discontinue | Money and people actually move | A comparison group, and an owner who can say stop |
Nunno puts the first one concretely: "Functional leaders who track every dollar spent by outcome category, such as productivity, revenue growth, risk mitigation, or innovation, are best positioned to defend investments and reallocate quickly if their projects underperform."
Outcome categories are not evenly loaded, which is why the category matters
Gartner found productivity was a key target outcome for 75% of functional leaders, and that it "commands approximately 30% of functional AI spend on average, nearly twice the percentage of the next highest objective." So the largest single block of AI spend is pointed at the outcome that is hardest to observe in a financial system. Revenue shows up in the ledger. Productivity has to be measured deliberately or it does not get measured at all.
Gartner also found the most popular use cases are not the highest returning ones. For IT, the most frequently pursued were cybersecurity threat detection and response (54%), IT service desk automation (54%) and automated code generation and refactoring (44%). The use cases where the largest proportion of leaders reported positive returns were intelligent IT asset and cost optimization (40%), synthetic data generation (28%) and automated code generation and refactoring (23%). Two separate questions with two separate denominators, so the figures cannot be subtracted. The point is the ordering, not a gap.
Almost nobody runs the fourth practice
Here is the number that reframes the whole discussion, and it is a failure number about the leaders themselves rather than a flattering comparison.
PwC's AI performance study, fielded in October and November 2025 across 1,217 director-level and above executives, found that while AI leaders are more disciplined than peers about pruning initiatives, "only 28% say they conduct AI portfolio reviews to terminate initiatives to a 'large' or 'very large' extent."
That is the top of the distribution, not the average organization. Read the scope honestly: 91% of that sample is publicly listed and 76% is at $1B or more in revenue, so it describes large, mostly listed companies sampled globally rather than smaller organizations.
Grant Thornton reaches the same practice from a different sample. Between February 23 and March 18 2026 it surveyed 950 US business leaders across 10 industries, drawn from senior leadership including CEOs, CFOs, COOs and CIOs or CTOs plus leaders reporting directly to the C-suite. Sumeet Mahajan, lead partner, AI and Data for Advisory Services, framed the requirement as needing to "apply discipline" and then to "set measurement targets, build governance infrastructure and curtail initiatives that do not deliver results."
Stopping things is the practice everyone endorses and nobody schedules. Continuing an underperforming initiative costs no one anything this quarter. Stopping it is a visible decision with a name attached to it.
The objection, from inside the same research
Worth carrying, because the strongest argument against a governance prescription comes from the firm making it. Tom Puthiyamadam, managing partner of advisory services at Grant Thornton, warns that "centralized review bodies become overwhelmed, creating bottlenecks that slow execution without reducing risk."
He is right, and it is the failure mode to design against. His own answer, in the same breath, is to "set policy and risk criteria centrally, then delegate assessments to trained reviewers at the division or regional level, aligning the depth of review to the level of risk." A review cadence that cannot keep pace with the initiative count becomes a queue, and a queue adds delay without adding judgment.
There is a second objection and it is stronger than the first. Gartner did not find that these four practices produce returns. It defined high performers as the companies running them, then measured what those companies reported. The practices and the group are the same thing, so the 81% describes the group rather than testing the practice. The direction could also run backwards: a company already getting returns can afford a review board and a spend taxonomy, while a company burning money on stalled pilots is cutting exactly those roles.
Nothing in these three studies separates those readings, and all three publishers sell advisory services in this area. Grant Thornton states plainly that its reported correlations are not guarantees of future results and should not be read as proof that one thing produced the other. What survives is narrower and still useful. Three independent samples across three fielding windows describe the same practice as present among the companies doing well and thin among the rest.
What makes a stop decision possible
The reason the fourth practice is rare is not mostly courage. Most organizations cannot support the sentence "this initiative underperformed."
Underperformed against what? A stop decision needs a comparison, and most AI reporting produces a single number with nothing to compare it to. Usage is up. Satisfaction is high. Self-reported time saved is meaningful. None of those can be underperformance, because none of them has anything to be measured against.
That is the gap our measurement methodology is built for: controlled cohort analysis with confidence scoring, comparing adopting teams to non-adopting teams while controlling for tenure, role and tool stack, with the limits disclosed alongside the result. Smaller claims, and claims that can carry a decision to stop funding something.
Two prerequisites sit underneath it, and both are unglamorous.
- Spend visibility at the level the review runs on. Gartner found roughly 11% of organizations are entirely unaware of what their function spent on AI in 2025. A portfolio review over spend nobody can total is theatre. SaaS Stack Intelligence covers the tool-level layer, utilization and duplicates and cost per active user. A function's total AI budget still has to come from finance.
- An outcome measure that is not an adoption count. Usage rises whether or not the work changed, which is why the AI Impact signal family treats adoption as the exposure variable in a controlled cohort comparison and surfaces high adoption with flat productivity as an enablement gap rather than a win.
Frequently asked questions
What is AI portfolio management?
Treating AI initiatives the way a portfolio manager treats holdings rather than as a list of projects that each run to completion. Gartner's definition contains four practices: continuous tracking, portfolio framing, scheduled reassessment, and reallocation or discontinuation.
Do companies that fail to track AI ROI lose visibility on their returns?
Gartner's release does not say that. It defines high performers by four practices, and separately reports what low performers said. It never defines low performers as the companies that fail to track, so the two figures cannot be joined into one claim.
How many organizations actually review AI initiatives in order to stop them?
Among the leaders, 28%, per PwC's study of 1,217 director-level and above executives fielded in October and November 2025. That is the top of the distribution, and the sample is large listed companies rather than smaller organizations.
Why is it hard to discontinue an AI initiative?
Partly organizational, because stopping is a visible decision and continuing is not. Mostly measurement, because you cannot say an initiative underperformed without something to compare it against.
What does an organization need in place before portfolio reviews are worth running?
Spend visibility at the level the review operates on, an outcome measure that is not an adoption count, and a comparison group so underperformance is a measured statement rather than an impression.
Where to start
Not with a review board. Pick one initiative, name the outcome it is supposed to move, and identify the group it should be compared against. If you cannot name the comparison group, the review would not have produced a decision anyway.
Request a Demo to see how outcome measurement and adoption signals are kept separate, or start a 90-Day AI Impact Audit.
Levos is accepting design partner applications from US organizations of 150 employees or more with an active AI rollout. Large organizations typically start with one function or division, measured against comparable teams that have not adopted yet. That is a stronger attribution design than a company-wide before-and-after.