Skip to Content

AI Usage in Performance Reviews: Why the Token Counter Broke as a Metric

September 14, 2026 • Levos Marketing • 7 min read
AI Usage in Performance Reviews: Why the Token Counter Broke as a Metric

AI usage counts measure consumption, not contribution. They belong on the input side of an impact model and cannot carry an outcome claim on their own. That has been an argument for two years. As of early September 2026 it is a documented case: one of the largest technology employers in the world put AI usage into its performance review guidance for roughly 10 months, then took that criterion back out.

Two things are true at once. Encouraging people to use a new tool is reasonable, and measuring whether they do is reasonable. Treating that number as evidence of impact is a different move, and it is the one that failed.

What actually happened, with dates

In a November 13 2025 internal memo reported by Business Insider, Meta's head of people Janelle Gale told employees that "AI-driven impact" would become a "core expectation" from 2026. Individual usage and adoption metrics were kept out of the 2025 annual cycle, so the change was prospective.

By April 2026 the input side had a scoreboard. Fortune, on a story broken by The Information, described an employee-built internal leaderboard called Claudeonomics, ranking the top 250 token users across a workforce Fortune puts at more than 85,000 people, with titles including "Token Legend" and "Cache Wizard." Over a 30-day window, usage tracked on that dashboard exceeded 60 trillion tokens, and the highest-ranked individual averaged 281 billion. Fortune estimated that one person's consumption at more than $1.4M using published list prices. The leaderboard came down two days after the story broke, and Meta told Fortune that "the employee took down the dashboard at their discretion; Meta did not request this action." WIRED reports the company began rationing employee AI usage a couple of months later.

Then, in an internal announcement the week of September 1 2026, the criterion came out. WIRED, which spoke to three employees who received the message, reported that the new guidance "replaces references to evaluating employees based on criteria like 'usage of AI' and their 'AI Native' designation, with looser wording, which caveats that 'these outcomes can be supported by AI or other means.'"

Meta's position is that this changes less than it appears. A spokesperson told WIRED the update emphasizes what was always the case, that Meta evaluates employees on their contributions, and that labels such as "AI Native" were never used for performance evaluation. The argument below does not depend on which account is right. What a usage count can support is the same either way.

Why adoption volume cannot carry an impact claim

Three defects, and they compound.

It is gameable by construction

Any metric that counts volume can be satisfied by producing volume. This is not a character observation about employees, it is a property of the measure. A leaderboard ranking people by tokens consumed rewards the behavior it was built to detect, which is what makes it useless as evidence.

The management literature described this long before AI. Ordóñez, Schweitzer, Galinsky and Bazerman set it out in "Goals Gone Wild" in 2009: a narrow goal concentrates effort on the goaled dimension and degrades attention to everything not goaled. An agentic tool makes that dimension cheaper to satisfy than ever, because a process left running overnight generates activity with nobody in the loop.

It has no denominator, so absence looks like underperformance

An activity metric counts what happened. It does not know what should have happened, or who was there to make it happen. Run it across a population without controlling for exposure and anyone who was not present scores lower for a reason that has nothing to do with their work.

This is the substance of a lawsuit filed in July 2026 in federal court in California, in which 26 current and former Meta employees allege that AI-assisted workplace analytics and productivity metrics, including dashboards tracking AI adoption and token usage, disadvantaged workers on legally protected medical, disability, family or parental leave during the company's layoffs. These are allegations and they are contested. A Meta spokesperson told CNBC the claims "lack merit and are not based on facts" and that "workforce management and organizational decisions were and are made by people, not AI." The case is ongoing.

A court will decide the legal question. The structural point stands either way: a metric with no denominator cannot tell a low score from a low opportunity to score.

Availability is not the only bias such a score imports. Writing in Forbes in November 2025, law professor Michelle Travis flagged a different exposure in the same design, citing research in which evaluators rated identical work as less competent when told AI had assisted it, with a materially larger penalty for women than men. Different mechanism, same lesson. Scoring people on AI use pulls in effects unrelated to the work.

The narrow fix is a denominator. The wide fix is a comparison group.

Hours-available as a denominator fixes the leave problem. It does not fix the harder one: you still have no idea what the usage bought. For that you need something to compare against.

The unit of measurement moves while you are measuring

WIRED reported that in the same week the criterion was removed, employees said their token consumption was continuing to surge, while the company was encouraging staff to test Hatch, an agentic tool that workers say consumes more than standard chatbots and coding assistants.

Look at what that does to the number. The incentive to inflate it was removed, rationing was already in place, and a tool that consumes more tokens per unit of work arrived. The line kept going up anyway.

You cannot read that curve. It is consistent with the tools becoming genuinely more useful, which is the strongest case here and should not be waved away. It is also consistent with the same work now costing more tokens. Nothing in the number tells you which, and the observation is employee self-report over days rather than a measured window. A token is a unit of model consumption, not a unit of output, and its relationship to output changes every time the tool changes.

Anyone tracking AI adoption through raw consumption is measuring with a ruler whose markings move whenever a vendor ships a release.

Input metrics and outcome metrics are not interchangeable

The fix is not to stop measuring adoption. It is a real and useful signal, and it sits inside the AI Impact family, one of the six signal families we track. The fix is to stop asking an input metric to do an outcome metric's job.

  Input metric Outcome metric
Example Tokens consumed, seats active, prompts sent Cycle time, throughput, quality, revenue per employee
What it answers Is the tool being used Did the work change
Fails when Volume is rewarded, or the tool changes Nothing to compare against
Safe use Coverage, rollout health, cost forecasting Impact claims, investment decisions
Unsafe use Individual performance ratings Reported without a stated confidence level

The dividing line is comparison. An input metric describes one group at one moment. An outcome metric only means anything against a group that did not get the treatment.

That is why our position on measurement methodology is controlled cohort analysis with confidence scoring rather than a single headline number. Adopting teams are compared to non-adopting teams while controlling for tenure, role and tool stack, and the limits are disclosed alongside the result. Smaller, less quotable claims that survive a finance review and a legal review. That is the trade.

Two design rules follow, and both are cheap.

  • Adoption signals inform the model. They never carry a rating on their own. The moment a usage count becomes the thing a rating is set against, it stops being a measurement and becomes a target, and the number degrades from that day forward.
  • Aggregate before you interpret. Individual data flows to a person's direct manager and not routinely up the chain, with a logged escalation exception for urgent cases. Team-level views require 5 or more people. That is our operating model on data handling, and it is a measurement rule as much as a privacy one: an n of 1 has no error bar.

The wider gap this sits inside

None of this is specific to one employer. Aon's 2026 Human Capital Trends Study, fielded November 2025 to January 2026 across 62 geographies with 2,361 board directors and senior business and people leaders, found 73% of organizations have deployed or are piloting AI programs while 18% report that most of their workforce participated in AI reskilling or upskilling in the past year.

Aon is documenting a readiness gap, not a measurement gap, and that distinction is worth keeping. But the shape is the same. Deployment is running ahead of the apparatus around it, and measurement is part of that apparatus. When a rollout arrives before the instrumentation, the first available number gets promoted into the role of evidence. Usage counts are available. That is most of why they get used.

Frequently asked questions

Should AI usage be part of performance reviews at all?

As context for a conversation, it is defensible. As a scored input to a rating, it is not. A usage count cannot separate productive use from volume generated to satisfy the metric, and it penalizes anyone whose availability was reduced for a protected reason.

What is the difference between an AI input metric and an AI outcome metric?

An input metric measures whether a tool is being used, for example tokens consumed or seats active. An outcome metric measures whether the work changed, for example cycle time or throughput. Input metrics suit rollout health and cost forecasting. Only outcome metrics, measured against a comparison group, support an impact claim.

Why are token counts unreliable as a productivity measure?

A token is a unit of model consumption, not a unit of output, and the tokens required for a fixed amount of work change whenever the tool changes. Agentic tools consume far more tokens per task than chat interfaces, so a rising token line can reflect a tool change rather than more work getting done.

How can an organization measure AI impact defensibly?

Compare adopting teams to non-adopting teams on outcome metrics, control for tenure, role and tool stack, attach a confidence level, disclose the limits, and keep adoption signals out of ratings.

Does removing usage from reviews mean AI adoption stopped mattering?

No. It means the measurement was assigned to the wrong variable. Adoption still matters as a leading indicator of coverage and cost. It is not evidence of return on its own.

Where this leaves you

If you are building an AI adoption dashboard, the useful question is not how to make the counter more accurate. It is which decisions the counter is allowed to touch.

That constraint is not a limitation on the measurement. It is the measurement.

Request a Demo to see how the AI Impact signal family separates adoption signals from outcome measurement, or start a 90-Day AI Impact Audit.

Levos is accepting design partner applications from US organizations of 150 employees or more with an active AI rollout. Large organizations typically start with one function or division, measured against comparable teams that have not adopted yet. That is a stronger attribution design than a company-wide before-and-after.

Business Insider, "Meta is about to start grading workers on their AI skills," Jyoti Mann, November 14 2025. Reports an internal memo from Janelle Gale dated November 13 2025. https://www.businessinsider.com/meta-ai-employee-performance-review-overhaul-2025-11

Fortune, "A Meta employee created a dashboard so coworkers can compete to be the company's No. 1 AI token user," April 9 2026. https://fortune.com/2026/04/09/meta-killed-employee-ai-token-dashboard/

WIRED, "Meta Pushes Its New AI Agent on Employees, but Eases Off on Tokenmaxxing," Paresh Dave and Maxwell Zeff, September 2 2026. Sourced to three employees who received the internal announcement. https://www.wired.com/story/meta-pushes-its-new-ai-agent-on-employees-but-eases-off-on-tokenmaxxing/

The Information, "Exclusive: Meta Tells Engineers AI Token Usage Won't Be Part Of Performance Reviews," September 2 2026. Broke the story. https://www.theinformation.com/briefings/exclusive-meta-tells-engineers-ai-token-usage-part-performance-reviews

eWeek, "Meta Faces Lawsuit Alleging AI Penalized Workers on Protected Leave," Joseph Chisom Ofonagoro, July 15 2026, citing Courthouse News Service for the complaint and CNBC for Meta's response. Allegations are contested and the case is ongoing. https://www.eweek.com/news/meta-ai-assisted-layoff-lawsuit-2026/

Forbes, "Meta's Plan To Evaluate Employees' AI Use May Negatively Impact Women," Michelle Travis, Research Professor at the University of San Francisco School of Law, November 20 2025. Cited for the evaluator-perception research she summarizes, not for the availability argument. https://www.forbes.com/sites/michelletravis/2025/11/20/metas-plan-to-evaluate-employees-ai-use-may-negatively-impact-women/

Aon, 2026 Human Capital Trends Study. Fielded November 2025 to January 2026, 2,361 board directors and senior business and people leaders across 62 geographies, released April 28 2026. Study and methodology: https://www.aon.com/en/insights/reports/human-capital-trends-study The 73% and 18% figures appear in the accompanying release: https://aon.mediaroom.com/2026-04-28-Nearly-90-percent-of-companies-believe-people-will-determine-AI-success,-but-far-fewer-are-investing-in-related-people-strategies,-Inaugural-Aon-Study-Finds

Ordóñez, Lisa D., Schweitzer, Maurice E., Galinsky, Adam D., and Bazerman, Max H. "Goals Gone Wild: The Systematic Side Effects of Overprescribing Goal Setting." Academy of Management Perspectives 23, no. 1 (2009): 6 to 16. Full text: https://www.hbs.edu/ris/Publication%20Files/09-083.pdf

Levos, "AI Impact." https://levos.ai/ai-impact

Levos, "Measurement Methodology." https://levos.ai/measurement

Levos, "How We Operate." https://levos.ai/how-we-operate

Share this article

Help others discover workforce intelligence insights

Levos Editorial

Levos Editorial publishes operator-grade research on workforce intelligence, AI deployment measurement, and human capital optimization. Reach the team at marketing@levos.ai