Skip to content
← Back to blog
·8 min read

Tokenomics: Nobody Knows Where the AI Money Goes

aimanagementoperations
Tokenomics: Nobody Knows Where the AI Money Goes

In April, Uber's CTO said the company had already spent its entire 2026 budget for AI coding tools. Four months into the year. In June, Bloomberg reported the fix: a cap of $1,500 a month per employee, per tool, with a dashboard to track it and an approval process to go over.

A few weeks before that, Microsoft told engineers in its Experiences and Devices division to stop using Claude Code by June 30 and move to GitHub Copilot CLI. The official reason was toolchain unification. Most of the reporting pointed at cost.

These are two of the most technically sophisticated buyers on the planet. Neither forecast it.

I don't read this as "AI is too expensive." I read it as something more boring and more dangerous. Companies are buying a product they can't measure. A KPMG survey reported by the Wall Street Journal found that only 26% of large US companies have a full, real-time view of what AI costs them. The rest see part of it, or see it when the invoice lands.

That's the real tokenomics problem. Not the price of a token. The fact that nobody inside the company can say where the tokens went.

The meter suits the seller

For a couple of years, most enterprise AI was sold like any other SaaS. A seat, a monthly fee, use it as much as you like within reason. That model died this year, quietly, one renewal at a time.

  • OpenAI aligned Codex pricing with API token rates on April 2, and extended it to every existing ChatGPT Enterprise plan on April 23.
  • Anthropic moved enterprise customers to a lower seat fee plus usage at API rates, applied as contracts come up for renewal.
  • GitHub moved every Copilot plan to usage-based billing on June 1, priced on input, output and cached tokens.

Simon Willison pieced together the timeline and his read is that both labs found product-market fit for coding agents. I agree with him. People overspend on things they actually want.

But look at why the meter is convenient for the vendor. A seat price is a bet on how much the average user will consume. Chat users were predictable. Agents aren't. An agent loops, retries, calls tools, spawns subagents and reads the same files twenty times. One engineer running agents unattended can burn more in a night than a chat user does in a month. With a flat seat, the vendor eats that variance. With a meter, you do.

That's a fair trade if you can see the meter. Most companies can't.

Cheaper tokens, bigger bills

The standard reassurance is that tokens keep getting cheaper, and they do. Stanford's 2025 AI Index measured the price of GPT-3.5-level performance falling from about $20 per million tokens in November 2022 to $0.07 in October 2024. A 280x drop for a fixed level of capability.

Nobody buys a fixed level of capability, though. You buy the newest model and you give it bigger jobs. In August, Gartner predicted that inference cost per agentic workflow will rise more than fivefold through 2028 and called it the inference paradox: cheaper tokens fund more complex workflows, which eat more tokens. It's a forecast, not a measurement. Anyone running agents has already seen the shape of it on an invoice.

So "it'll get cheaper" isn't a cost strategy. It's a hope.

I can't see it either

I'd love to say this is a big-company problem. It isn't.

When we built the autopilot for my own project (Part 3 of that series has the details), one of the stop rules was a spending cap per work item. On the day we built it, the cost meter recorded nothing usable for the work itself. One agent session held several items at once, so the spend couldn't be attributed to any of them. One repo, one person, and I still couldn't say where the money went.

Even when attribution worked, the number lied. A feature that cost $153.65 to build caused two fixes later. Its true cost was $219.78.

And when I finally sat down and read my own ledger, it was worse than I expected. About $7,270 of tracked agent spend, every priced call on a top-tier model, and 39% of it on phases I had labeled mechanical myself. I wrote that one up in Stop Paying Frontier Prices for If-Statements.

Now multiply that by forty teams, five tools, three model providers and a few hundred engineers.

The State of FinOps 2026 report says 98% of FinOps practitioners now manage AI spend, up from 31% in 2024. Their top challenges are the ones I hit: visibility into AI costs, allocating them to the business units that caused them, and proving value. The practice exists. The plumbing mostly doesn't, yet.

The gap is the next market

This is the part I find most interesting. Every time a big line item moves from a fixed price to a meter, an industry shows up to read the meter. Cloud did this a decade ago. Nobody understood their AWS bill, and a whole category of FinOps tools grew up to explain it, allocate it and argue about it.

Tokens are going through the same thing, faster.

Engineering intelligence platforms are adding spend. Jellyfish now tracks token usage and cost by team, developer, tool and model, plotted against merged pull requests. Their own research on Q1 2026 data from about 12,000 developers shows why that matters. The heaviest 20% of users spent about $1,822 in tokens for the quarter and averaged 23 merged PRs. The lightest 20% spent about $3 and averaged 11. Roughly twice the output for about 600 times the spend. That's vendor data, priced at list API rates rather than real invoices, so treat it as directional. But it's exactly the question a CFO will ask, and today most engineering orgs couldn't answer it.

The big platforms are buying in. Atlassian paid about $1 billion for DX last year, and pitched the deal as helping customers work out whether their AI investments are paying off.

Standards are forming. In June the Linux Foundation announced the Tokenomics Foundation, backed by Google Cloud, Microsoft, IBM, JPMorganChase, KPMG and others. It will work with the FinOps Foundation to extend the FOCUS billing spec to token-based spend. Jim Zemlin put it simply: "tokens have become the new unit of technology spend." When the Linux Foundation creates a foundation around your cost line, it's no longer an engineering footnote.

Add the LLM gateways and proxies that sit between your code and the model APIs, and you can see the stack forming. Capture at the gateway, attribute in the engineering data, report in the FinOps tool.

What the dashboards won't fix

A dashboard tells you who spent what. It doesn't tell you whether it was worth it.

Merged PRs are a weak proxy for value. So are lines of code and ticket counts. Jellyfish says as much about its own token metric: useful, directionally, and incomplete. Pick the wrong unit and you'll optimize the wrong thing with great precision.

Visibility doesn't fix incentives either. Uber reportedly ran internal leaderboards ranking teams by AI usage. Reward consumption and you get consumption. The cap was a fix for a metric the company chose.

And caps hit the wrong people. A flat limit lands hardest on your heaviest users, and some of them are your most productive engineers. The goal isn't to spend less. It's to know what each dollar bought.

Who owns the meter

My take: this belongs to engineering leadership, not finance. Finance owns the budget. But the decisions that drive the bill are engineering decisions. Which model. Which tool. How agents are allowed to run and when they stop. If engineering doesn't own the meter, finance will own it with a blunt instrument, and the blunt instrument is always a cap.

What to do Monday morning

  1. Find every meter. List every place the company buys tokens: API keys, coding tools, SaaS products that quietly switched to usage pricing. Check the renewal dates, because that's where the price model changes.
  2. Route model calls through one gateway. Make team, product and workload tags mandatory. Untagged spend is the problem, not a rounding error.
  3. Pick one unit metric per workload. Cost per merged PR, per resolved ticket, per document processed. Imperfect is fine. None is not.
  4. Give unattended agents stop rules. A cost cap per item, a failure limit, a kill switch. I wrote about ours in The Driver Is a Script, Not an Agent.
  5. Kill usage leaderboards. Replace them with output per dollar. And get the visibility layer in place, bought or built, before your next renewal, not after it.

Tokens will keep getting cheaper. Your bill won't, unless somebody can read it.