AI usage counts measure consumption, not contribution. They belong on the input side of an impact model and cannot carry an outcome claim on their own. That has been an argument for two years. As of early September 2026 it is a documented case: one of the largest technology employers in the world put AI usage into its performance review guidance for roughly 10 months, then took that criterion back out.
Two things are true at once. Encouraging people to use a new tool is reasonable, and measuring whether they do is reasonable. Treating that number as evidence of impact is a different move, and it is the one that failed.
What actually happened, with dates
In a November 13 2025 internal memo reported by Business Insider, Meta's head of people Janelle Gale told employees that "AI-driven impact" would become a "core expectation" from 2026. Individual usage and adoption metrics were kept out of the 2025 annual cycle, so the change was prospective.
By April 2026 the input side had a scoreboard. Fortune, on a story broken by The Information, described an employee-built internal leaderboard called Claudeonomics, ranking the top 250 token users across a workforce Fortune puts at more than 85,000 people, with titles including "Token Legend" and "Cache Wizard." Over a 30-day window, usage tracked on that dashboard exceeded 60 trillion tokens, and the highest-ranked individual averaged 281 billion. Fortune estimated that one person's consumption at more than $1.4M using published list prices. The leaderboard came down two days after the story broke, and Meta told Fortune that "the employee took down the dashboard at their discretion; Meta did not request this action." WIRED reports the company began rationing employee AI usage a couple of months later.
Then, in an internal announcement the week of September 1 2026, the criterion came out. WIRED, which spoke to three employees who received the message, reported that the new guidance "replaces references to evaluating employees based on criteria like 'usage of AI' and their 'AI Native' designation, with looser wording, which caveats that 'these outcomes can be supported by AI or other means.'"
Meta's position is that this changes less than it appears. A spokesperson told WIRED the update emphasizes what was always the case, that Meta evaluates employees on their contributions, and that labels such as "AI Native" were never used for performance evaluation. The argument below does not depend on which account is right. What a usage count can support is the same either way.
Why adoption volume cannot carry an impact claim
Three defects, and they compound.
It is gameable by construction
Any metric that counts volume can be satisfied by producing volume. This is not a character observation about employees, it is a property of the measure. A leaderboard ranking people by tokens consumed rewards the behavior it was built to detect, which is what makes it useless as evidence.
The management literature described this long before AI. Ordóñez, Schweitzer, Galinsky and Bazerman set it out in "Goals Gone Wild" in 2009: a narrow goal concentrates effort on the goaled dimension and degrades attention to everything not goaled. An agentic tool makes that dimension cheaper to satisfy than ever, because a process left running overnight generates activity with nobody in the loop.
It has no denominator, so absence looks like underperformance
An activity metric counts what happened. It does not know what should have happened, or who was there to make it happen. Run it across a population without controlling for exposure and anyone who was not present scores lower for a reason that has nothing to do with their work.
This is the substance of a lawsuit filed in July 2026 in federal court in California, in which 26 current and former Meta employees allege that AI-assisted workplace analytics and productivity metrics, including dashboards tracking AI adoption and token usage, disadvantaged workers on legally protected medical, disability, family or parental leave during the company's layoffs. These are allegations and they are contested. A Meta spokesperson told CNBC the claims "lack merit and are not based on facts" and that "workforce management and organizational decisions were and are made by people, not AI." The case is ongoing.
A court will decide the legal question. The structural point stands either way: a metric with no denominator cannot tell a low score from a low opportunity to score.
Availability is not the only bias such a score imports. Writing in Forbes in November 2025, law professor Michelle Travis flagged a different exposure in the same design, citing research in which evaluators rated identical work as less competent when told AI had assisted it, with a materially larger penalty for women than men. Different mechanism, same lesson. Scoring people on AI use pulls in effects unrelated to the work.
The narrow fix is a denominator. The wide fix is a comparison group.
Hours-available as a denominator fixes the leave problem. It does not fix the harder one: you still have no idea what the usage bought. For that you need something to compare against.
The unit of measurement moves while you are measuring
WIRED reported that in the same week the criterion was removed, employees said their token consumption was continuing to surge, while the company was encouraging staff to test Hatch, an agentic tool that workers say consumes more than standard chatbots and coding assistants.
Look at what that does to the number. The incentive to inflate it was removed, rationing was already in place, and a tool that consumes more tokens per unit of work arrived. The line kept going up anyway.
You cannot read that curve. It is consistent with the tools becoming genuinely more useful, which is the strongest case here and should not be waved away. It is also consistent with the same work now costing more tokens. Nothing in the number tells you which, and the observation is employee self-report over days rather than a measured window. A token is a unit of model consumption, not a unit of output, and its relationship to output changes every time the tool changes.
Anyone tracking AI adoption through raw consumption is measuring with a ruler whose markings move whenever a vendor ships a release.
Input metrics and outcome metrics are not interchangeable
The fix is not to stop measuring adoption. It is a real and useful signal, and it sits inside the AI Impact family, one of the six signal families we track. The fix is to stop asking an input metric to do an outcome metric's job.
| Input metric | Outcome metric | |
|---|---|---|
| Example | Tokens consumed, seats active, prompts sent | Cycle time, throughput, quality, revenue per employee |
| What it answers | Is the tool being used | Did the work change |
| Fails when | Volume is rewarded, or the tool changes | Nothing to compare against |
| Safe use | Coverage, rollout health, cost forecasting | Impact claims, investment decisions |
| Unsafe use | Individual performance ratings | Reported without a stated confidence level |
The dividing line is comparison. An input metric describes one group at one moment. An outcome metric only means anything against a group that did not get the treatment.
That is why our position on measurement methodology is controlled cohort analysis with confidence scoring rather than a single headline number. Adopting teams are compared to non-adopting teams while controlling for tenure, role and tool stack, and the limits are disclosed alongside the result. Smaller, less quotable claims that survive a finance review and a legal review. That is the trade.
Two design rules follow, and both are cheap.
- Adoption signals inform the model. They never carry a rating on their own. The moment a usage count becomes the thing a rating is set against, it stops being a measurement and becomes a target, and the number degrades from that day forward.
- Aggregate before you interpret. Individual data flows to a person's direct manager and not routinely up the chain, with a logged escalation exception for urgent cases. Team-level views require 5 or more people. That is our operating model on data handling, and it is a measurement rule as much as a privacy one: an n of 1 has no error bar.
The wider gap this sits inside
None of this is specific to one employer. Aon's 2026 Human Capital Trends Study, fielded November 2025 to January 2026 across 62 geographies with 2,361 board directors and senior business and people leaders, found 73% of organizations have deployed or are piloting AI programs while 18% report that most of their workforce participated in AI reskilling or upskilling in the past year.
Aon is documenting a readiness gap, not a measurement gap, and that distinction is worth keeping. But the shape is the same. Deployment is running ahead of the apparatus around it, and measurement is part of that apparatus. When a rollout arrives before the instrumentation, the first available number gets promoted into the role of evidence. Usage counts are available. That is most of why they get used.
Frequently asked questions
Should AI usage be part of performance reviews at all?
As context for a conversation, it is defensible. As a scored input to a rating, it is not. A usage count cannot separate productive use from volume generated to satisfy the metric, and it penalizes anyone whose availability was reduced for a protected reason.
What is the difference between an AI input metric and an AI outcome metric?
An input metric measures whether a tool is being used, for example tokens consumed or seats active. An outcome metric measures whether the work changed, for example cycle time or throughput. Input metrics suit rollout health and cost forecasting. Only outcome metrics, measured against a comparison group, support an impact claim.
Why are token counts unreliable as a productivity measure?
A token is a unit of model consumption, not a unit of output, and the tokens required for a fixed amount of work change whenever the tool changes. Agentic tools consume far more tokens per task than chat interfaces, so a rising token line can reflect a tool change rather than more work getting done.
How can an organization measure AI impact defensibly?
Compare adopting teams to non-adopting teams on outcome metrics, control for tenure, role and tool stack, attach a confidence level, disclose the limits, and keep adoption signals out of ratings.
Does removing usage from reviews mean AI adoption stopped mattering?
No. It means the measurement was assigned to the wrong variable. Adoption still matters as a leading indicator of coverage and cost. It is not evidence of return on its own.
Where this leaves you
If you are building an AI adoption dashboard, the useful question is not how to make the counter more accurate. It is which decisions the counter is allowed to touch.
That constraint is not a limitation on the measurement. It is the measurement.
Request a Demo to see how the AI Impact signal family separates adoption signals from outcome measurement, or start a 90-Day AI Impact Audit.
Levos is accepting design partner applications from US organizations of 150 employees or more with an active AI rollout. Large organizations typically start with one function or division, measured against comparable teams that have not adopted yet. That is a stronger attribution design than a company-wide before-and-after.