Why AI Output Metrics Mislead and What to Track Instead
Table of contents
- Why AI Output Metrics Mislead and What to Track Instead
- The Metrics We Default To
- When Output Metrics Lie
- The Counter-Signal Nobody Expected
- What the Dashboard Cannot See
- Three Questions That Change the Frame
- The Cost the Dashboard Hides
- Innovation Accounting: The Closest Precedent
- The Objection Worth Taking Seriously
- Where to Start
- The Measurement That Matters
Why AI Output Metrics Mislead and What to Track Instead
Sixteen experienced software developers sat down to work. They had access to the best AI coding tools available. They had been told, reasonably, to expect a 20 to 25 percent productivity boost. When the study ended, they believed, on average, that AI had made them 20 percent faster. They were 19 percent slower. That is from METR’s July 2025 randomized controlled trial, the most rigorous study of AI’s effect on experienced developer productivity published to date. The 39-point gap between what developers believed about their performance and what actually happened is not a rounding error. It is a measurement failure at scale. And it should be the first thing any executive reads before reviewing their organization’s AI productivity numbers.

The Metrics We Default To
Most organizations are measuring the wrong things. Not because their teams are careless, but because the wrong things are easy to count. Tokens consumed. Lines of code generated. Pull requests merged. Tool adoption rates. Story points completed. Time-to-first-response in customer service queues. Meeting transcripts enabled. These are the metrics that appear on AI dashboards across the enterprise right now. Every one of them produces a number. Every one of them trends over time. Every one of them can be presented in a board slide. None of them tell you whether your AI transformation is creating value. This is output accounting. It measures what happened: how much AI was used, how many tasks were touched, and how fast certain activities were completed. It does not measure whether the right things happened, whether the quality of work improved, or whether the organization can now do anything it could not do before. The appeal is understandable. Output metrics are fast, cheap, and unambiguous. In a transformation that feels uncertain and fast-moving, the comfort of a trending dashboard is real. That comfort is the problem.
When Output Metrics Lie
The clearest proof is Klarna. In 2024, Klarna announced that AI had replaced approximately 700 customer service agents. By Q3 2025, the company claimed its AI agent was doing the work of 853 full-time employees and saving $60 million annually. The volume metrics looked excellent: faster resolution times, higher tickets-per-hour throughput, and lower cost-per-interaction. Then, quietly, in early 2026, Klarna began rehiring humans. What the output metrics had not captured was quality deterioration on complex interactions. Customer satisfaction scores on difficult cases had declined. The conversations that mattered most, the ones where customers were frustrated and needed real understanding, were getting worse. The metrics that declared success had optimized for speed in the cases where speed was least important. Klarna’s reversal is not a story about AI failing. It is a story about measurement failing. The organization tracked what was easy to track. The things that were hard to track, judgment quality on complex cases, customer trust in sensitive interactions, brand perception over time, deteriorated while the dashboard numbers climbed. This is not unique to Klarna. It is the predictable outcome of any measurement system that optimizes for outputs rather than outcomes.
The Counter-Signal Nobody Expected
Triumph Financial runs one of the largest payment networks in trucking, moving roughly $18 billion a year to carriers who often wait up to 90 days to get paid by brokers and shippers. In 2023, working with KUNGFU.AI, the company set out to speed up invoice funding with AI. It would have been easy to make speed the headline metric. Triumph didn’t.
Earlier attempts at automation, using rigid, hard-coded rules, had already failed once, rejecting too many good invoices to be useful. So this time, before anyone touted a faster approval time, the team agreed on three specific numbers to watch: disputed invoices, short pays, and write-offs. Not throughput. Not approval speed. The things that would reveal whether the model’s decisions were actually as good as a human’s, not just faster than one. Those numbers were put on a shared dashboard everyone could see, and the model wasn’t scaled up until a staged rollout, moving from historical data testing to a dark launch to a 100-day pilot, showed it held up.
Only after that did the speed numbers get to matter. And they mattered a great deal: invoice approval time fell from an average of 178 minutes to about 10 seconds, with more than $4 billion in invoices now running through the model and over half auto-approved. But the figures Triumph’s CTO Jason Heilig points to first aren’t the speed ones. Short pay is down 57 percent. Chargebacks are down 65 percent. Disputes are down 25 percent. The company didn’t get fast at the expense of getting careless. It got fast because it insisted on being careful first.
That is the inverse of Klarna. Klarna’s dashboard celebrated speed and volume while the quality of the hardest conversations quietly eroded underneath it. Triumph made quality the metric that had to clear the bar before speed was allowed to become the story at all.
What the Dashboard Cannot See
Output metrics fail for a structural reason. They measure what was done. What leaders actually need to know is whether things are getting better. That distinction sounds simple. It produces completely different questions. Counting tokens consumed is easy. Asking whether the judgment underlying those tokens improved is hard. Counting PRs is easy. Asking whether the engineering organization is now capable of work it could not previously do is hard. Counting tool adoption is easy. Asking whether the team’s AI outputs are being accepted, revised, or rejected, and learning from the pattern, is hard. The organizations that stay in output-accounting mode past the early adoption phase are not being prudent. They are deferring accountability. McKinsey’s November 2025 “State of AI” report found that 88 percent of organizations now use AI, but only 6 percent qualify as high performers with measurable EBIT impact. Only 39 percent report any measurable business effect at all. The gap between having AI and benefiting from AI is the gap between output accounting and outcome accounting. High performers, in McKinsey’s data, are 2.8 times more likely to have fundamentally redesigned workflows. That is not a technology finding. It is a measurement finding: the organizations that ask deeper questions build different systems.
Three Questions That Change the Frame
Output metrics answer the question “did we use AI?” The questions leaders actually need are harder and closer to the truth. The first: how many agents do you have running? This measures the scale of actual deployment, not licenses activated or employees who opened a tool. Agents running against real problems are a more honest indicator of organizational AI maturity than any adoption metric. The second: how long can those agents run without human intervention? This is a proxy for quality. An agent that requires correction every five minutes signals weak prompting, poor context engineering, or an immature tool setup. An agent that completes substantive work over hours signals a team that has genuinely learned to work with AI. Autonomy duration is a quality metric dressed as a timing metric. Anthropic’s internal research found that human interventions per Claude Code session fell from 5.4 to 3.3 between August and December 2025\. That four-month trajectory is more revealing than any adoption curve. The third: what novel work are you unlocking that was not feasible before? This is the one that matters most to leaders, investors, and boards. Not “we are shipping the same roadmap 20 percent faster” but “we stood up a customer-segmentation pipeline that had been on the backlog for two years because we could never justify the engineering cost.” Novel work is opportunity creation. It is the metric that corresponds to what organizations are actually hoping for when they invest in AI. None of these three questions are easy to put on a dashboard. That is the feature, not the bug. A metric that is easy to optimize is a metric that will be gamed. Consider what happened in the rooms where CEOs set personal token-consumption targets for their teams: staff ran purposeless prompts to hit the number. Output metrics corrupt the behavior they are meant to measure. Outcome metrics resist that corruption because they are attached to something real.

The Cost the Dashboard Hides
There is a second blind spot, and it sits on the other side of the ledger. The three questions above measure whether the work is getting better. They do not measure what it now costs to deliver it, and that is the other half of any honest ROI. For two decades, software ran on an economic assumption so reliable that most leaders stopped noticing it. The marginal cost of serving one more user was effectively zero. Build the product once, and the ten-thousandth user cost almost nothing more than the thousandth. Engagement was therefore an unalloyed good. More usage meant more value, more retention, more expansion, and almost no additional cost to carry it. AI breaks that assumption. Every interaction now consumes tokens, and tokens cost money in direct proportion to use. Jeff Gothelf, who co-authored Lean UX, put the consequence plainly: your most engaged users can quietly become your least profitable ones. The power user running fifty AI queries a day is the user you celebrated under the old economics and the user who erodes your margin under the new one. The dashboard that shows engagement climbing may also be showing cost climbing faster, and a usage metric will never tell you which. This is why measurement for AI cannot stop at value. It has to track cost-to-serve at the unit level: cost per successful task, gross margin per active user, model cost as a percentage of revenue. These are not finance-team afterthoughts to reconcile at quarter end. They are leading indicators of whether an AI product or workflow stays economically sustainable as it scales. An organization can be creating genuine value, clearing every outcome bar in this piece, and still be quietly building something that gets less profitable with every new power user it celebrates. The discipline is the same one this entire piece argues for. Measure the thing that is hard to see, not the thing that is easy to count. On the value side, that means outcomes over outputs. On the cost side, it means cost per successful result over raw usage volume. A serious AI scorecard holds both, because a transformation that creates value while quietly destroying margin is not a success the dashboard is equipped to catch.
Innovation Accounting: The Closest Precedent
The intellectual framework that comes closest to what is needed already exists. Eric Ries built it for a different context: startups trying to measure progress when traditional indicators, revenue, customers, and profitability, are all effectively zero. His answer was innovation accounting. Instead of revenue, measure validated learning. Instead of units shipped, measure hypothesis tests completed. Build, measure, learn is the loop. The goal is not to produce a big number. The goal is to reduce uncertainty faster than the competition. Nobody has operationalized innovation, accounting for enterprise AI transformation in a published, replicable form. That gap is confirmed across every major measurement research program. DORA has extended its software-delivery metrics toward AI. Anthropic has published a primitives framework that measures how AI is being used at the task level. Accenture and Wharton have built a skills-shift index tracking 150 million professional profiles. None of them answer whether the organization is learning faster, producing better work, or doing things it could not do before. The practitioner community is ahead of the published literature on this. Six consecutive executive dinners across Dallas, Houston, Boston, Boulder, Portland, and Raleigh surfaced innovation accounting independently, without anyone being prompted. In every room, leaders described the same measurement problem and reached for the same frame. In no room did anyone have an operationalized version to point at. That convergence is not a coincidence. It means the field is ready for a framework, and the gap is structural, not a matter of individual companies being slow.
The Objection Worth Taking Seriously
The obvious counter-argument is that outcomes are unmeasurable. That output metrics are at least something, while outcome metrics are a vague ambition. This is a legitimate critique of poorly defined outcome goals. It is not a reason to abandon outcome measurement. Anthropic published the AI Fluency Index in early 2026, analyzing nearly 10,000 human-AI conversations to measure the quality of collaboration, not just the quantity. They identified 24 specific behaviors associated with effective AI use across four dimensions. Quality measurement is not an aspiration. It is an active research program producing real findings. The staged-measurement argument has something to it. Token consumption is a reasonable early proxy when the goal is normalizing AI use across a skeptical organization. Cultural adoption does need to come first. But most large organizations are past that phase now. Adoption is widespread. The question is no longer “will people use this?” It is “are we getting better because of it?” The Goodhart’s Law argument cuts both ways. Yes, any metric will eventually be gamed. That is a reason to rotate metrics deliberately, to pair quantitative measures with qualitative judgment, and to build measurement systems that are harder to optimize against than a single dashboard number. It is not a reason to accept measurements that are actively misleading. Outcomes can be defined. What novel work did your organization do this quarter that was not feasible last quarter? What percentage of AI outputs required human revision, rejection, or acceptance? How has autonomy duration changed over six months? These are measurable. They require judgment to interpret, as all meaningful metrics do. That is not a bug. That is what accountability looks like.
Where to Start
Auditing your current AI measurement against a single question produces clarity quickly: does each metric measure what happened or whether things got better? Tokens consumed, meeting transcripts enabled, PR counts, story points completed: these measure what happened. Replace them with counter-metrics that track quality alongside volume. If PR count goes up, track average PR complexity alongside it. If resolution time goes down, track customer satisfaction on complex interactions alongside it. At one roughly 1,000-person software company, the PR count told one story and average PR size told a different one. Both were necessary to understand what was actually happening. Then introduce the three questions as a leadership review practice. Not a dashboard, but a quarterly conversation: how many agents are running, how long can they run unattended, and what novel work did AI unlock in the last 90 days that was not on the roadmap before? The answers will be uneven and sometimes uncomfortable. That is the point. The organizations that build the discipline now, before their output metrics lock them into the wrong optimization, are the ones that will join the six percent achieving real EBIT impact. Not because outcome accounting is easy, but because it is honest.
The Measurement That Matters
Stanford HAI named this moment the shift from the era of AI evangelism to the era of AI evaluation. The question is not whether AI is transforming your industry. That question is answered. The question is whether your organization is learning from that transformation or simply counting it. Output accounting produces numbers that trend upward while the business-critical questions go unanswered. Klarna had great numbers until it did not. The engineering executive had a disappointing PR count until someone looked at PR size. The METR developers believed they were 20 percent faster until the data showed they were 19 percent slower. The dashboard that looks best right now may be hiding your Klarna moment. Outcome accounting is not optional for leaders who want to know what is actually happening. It is the discipline that makes the difference visible before it becomes irreversible. Want to explore what an outcome-accounting framework looks like for your organization? We work with leadership teams on exactly this question.