A mid-level leader at a technology company told me recently that two colleagues were each spending several thousand dollars a month on AI tools. Their usage numbers had become a talking point for the leadership team — held up as evidence the investment was working. Meanwhile, one of the company's top architects was spending a fraction of that, doing focused, precise work that directly moved a product forward. In the conversation he described, the architect's restraint was being read as resistance. The high spenders were the stars.
He wasn't sure how to say out loud what seemed obvious to him: that wise use and high use are not the same thing.
Around the same time, I was talking with an executive at a company that sells AI in its core product — a company, in other words, that has every reason to believe in what the technology can do. When I asked how they measured whether AI was working internally, the answer was immediate: the only real proof was a decrease in headcount. Nothing else counted.
Two organizations. Two different metrics. The same underlying problem.
When AI hit organizations at scale, the first question was whether it was real. Executives who decided yes — that this wasn't a passing moment — faced immediate pressure to justify the investment. The metrics they reached for weren't invented in a vacuum. The media, the labs, and investors were all pointing at the same first-order consequence: jobs at risk. Headcount became the proof point because the whole world had handed it to them. It was legible, reportable, and aligned with the dominant narrative. If AI was coming for work, then showing that it had reduced labor costs was the obvious way to demonstrate it was doing something.
The second wave came when leaders realized they needed people to actually use the tools rather than fear them. Activity became the signal. Usage counts, logins, spend — anything that showed people were engaging. This made sense as an entry point. You can't learn what AI changes about work without experimenting with it, and at the individual level that experimentation still matters. People develop real skills by doing. The problem was that exploration-phase metrics never got replaced when the exploration phase ended. Usage kept being reported as proof of progress long after it had stopped measuring anything useful.
Which brings most organizations to where they are now: headcount and activity as the gold standard, with no clear path from either to whether the work has actually gotten better.
Imagine a company in 1996 evaluating the productivity gains from email by measuring how many emails each employee sent per week. The top senders get recognized. The person who writes three precise messages that move decisions forward goes unnoticed. The metric captures activity. It tells you nothing about whether the communication produced anything.
Nobody would design that measurement system on purpose. But that's effectively what spend-as-proxy and headcount reduction are: the most available, countable things, adopted under pressure before anyone had time to ask what they were actually trying to learn.
Sales organizations figured this out a long time ago. You don't pay a rep on calls alone. Activity matters as a leading indicator, but the rep gets paid when they close. The activity metric exists to coach behavior. The outcome metric is what the organization is actually trying to produce. One is not a substitute for the other.
None of this is an indictment of the leaders who adopted these metrics. They were navigating genuine uncertainty under real pressure, with what was available to them at the time. The narrative that AI eliminates jobs wasn't invented by CFOs — it was handed to them fully formed. And asking people to experiment and then measuring whether they did was a reasonable response to a technology nobody fully understood yet.
The problem is that the moment has moved and the metrics haven't. Organizations still measuring adoption by activity are asking a three-year-old question. And organizations that can only demonstrate AI's value through headcount reduction have, perhaps without meaning to, defined away every other form of value it might create.
Better measures exist. None of them are easy to report on a slide, but all of them are closer to what organizations are actually trying to accomplish.
Revenue and pipeline impact. Can you connect AI-assisted work to deals closed, proposals accelerated, or customer problems resolved faster? This is the outcomes version of the sales analogy — not how many calls, but whether they closed.
Decision speed. How long does it take to move from a question to a commitment? If AI is working, that gap should be narrowing where it has been deployed intentionally. If it isn't, the tools may be running but the architecture around them hasn't changed.
Rework rate. How often does work get thrown away, relitigated, or sent back for revision? Rework is expensive, demoralizing, and largely invisible in most measurement systems. It's one of the most honest signals available: AI that improves clarity and alignment should reduce it.
Thinking time. Are your people doing work they couldn't do before — harder analysis, more considered decisions, problems that used to fall through the cracks — or are they just doing the same work faster? Recovered capacity is a leading indicator that the system is getting less wasteful. What people choose to do with that capacity is the real measure of whether any of this mattered.
These are harder to report than a headcount number or a usage dashboard. They require having agreed on what the organization is trying to accomplish with deploying the tools — which is the step organizations often skip. That's not a technology problem. It's the question that was always underneath the investment, waiting to be asked.