Cheaper tokens, bigger bills

AI keeps getting cheaper, the bills keep getting bigger, and the build-out still cannot earn itself back. Why falling token prices squeeze the companies that paid for the data centers instead of saving them.

When DeepSeek's cheap model rattled the AI trade in early 2025, Microsoft's chief executive Satya Nadella answered it in four words: "Jevons paradox strikes again." As AI gets cheaper, he wrote, its use will skyrocket into a commodity we cannot get enough of. Sam Altman says the cost to use a given level of AI falls about tenfold a year. Andreessen Horowitz named the trend "LLMflation": the price of a million tokens at a fixed quality has fallen roughly a thousandfold in three years. The reassurance underneath all of it is the same. AI keeps getting cheaper, so the trillion-dollar build-out will grow into itself, and everyone wins.

The first half is true. The conclusion is backwards. For the companies that paid for the build-out, cheaper tokens are not the escape from the cost problem. They are the trap.

Cheaper tokens, bigger bills

A token is the small chunk of text AI is billed by, and its price really has collapsed.

Top-tier output pricing fell about 4x since 2023; mid-tier quality fell 50 to 100x. Inference matching an older model's quality fell roughly 280x, from about $20 to $0.07 per million tokens, in two years. Andreessen Horowitz calls the trend LLMflation and clocks it at about 1,000x over three years.

The bull case rests on the Jevons paradox: the old observation that when something gets cheaper to use, people use so much more of it that total spending goes up, not down. Cheaper engines burned more coal, not less. Cheaper tokens, the argument goes, get used in such volume that revenue grows into the build-out.

And total spending is going up, because of what the cheaper tokens are being spent on. AI is shifting from chatbots, which answer one question, to agents, which plan a task, call tools, check their own work, and retry, burning many times the tokens to do it. Agentic work runs five to thirty times the tokens of a single chatbot answer, and the heaviest workflows over a thousand times more. So the price per token falls while the tokens per task explode faster, and the bill rises. Andreessen Horowitz's own figures show it: spending grew about 320 percent even as the per-token price fell 280-fold. Cheaper tokens, bigger bills.

That is good news for the company selling the compute, which is why NVIDIA's numbers look the way they do. It says nothing about whether the companies that bought all that compute can ever earn it back. They cannot, and here is why.

The trap

The Jevons argument assumes a falling price grows the seller's revenue without limit. For a company carrying hundreds of billions in data-center spending against a finite market, four things break that at once.

There are only so many buyers. You can sell your way out of a falling price only if demand grows forever. It does not. Maybe four billion people can realistically use this, and most company-wide AI rollouts still do not earn back what they cost. There is a ceiling on how many tokens there are to sell, and you cannot out-run a falling price once you hit it.

The build-out is a fixed bill. The hundreds of billions a year already committed to data centers, plus the cost of the chips wearing out and being written off, have to be paid no matter how cheap a token gets. Every cent shaved off the price is a cent less covering that bill, so paying it off takes ever more volume, which runs straight into the ceiling above.

The floor is set by the hardware you own. Once a token is cheap enough, customers notice they can run a capable free model on a computer they already bought, for almost nothing. The 2026 wave of AI laptops and desktops, covered in Arming both sides, makes the everyday, high-volume work runnable in-house. So the cheaper the cloud's tokens get, the stronger the case to stop renting and run it yourself. The efficiency driving the price down is the same force that lets the work walk off the cloud, onto hardware people own.

So the cloud is caught both ways. Drop the price to keep customers, and the build-out never gets paid off. Hold the price up to pay it off, and the everyday work walks to cheap local hardware. Either way, the trillion-dollar bet does not earn itself back at the cloud layer. Cheaper tokens are not the rescue. They are the squeeze.

Even at full scale, the math does not close

Grant the best case: the whole reachable market signs up. What does serving it earn, and what does it cost?

At today's per-user economics, serving roughly four billion people brings in about $340 to $500 billion a year. The data centers needed to serve that market cost on the order of $12 to $15 trillion to build, which is well over a trillion dollars a year just to cover them, and several trillion if the chips are written off over three years rather than ten, closer to how hard they are actually run. At every assumption, the yearly cost of serving the market is larger than the yearly revenue from serving it.

Morgan Stanley already puts the build-out near $3 trillion by 2028, and reckons the companies can cover only about half of it from their own cash; the rest gets borrowed. The price that would close the gap between what the market pays and what the build-out costs is far above where token prices are heading. It is not getting closer. It is getting further away, one price cut at a time.

The sellers know, and push the other way

"The market will make AI cheap for everyone" assumes the companies holding the bill are neutral. They are not. They have every reason to push the other way, and you can see it: steering everyone toward agentic workloads that burn more tokens, bundling and locking customers in so the bill does not shrink, fencing off the local and open path that would let the work escape, and closing the funding gap by quietly loading the public in through index funds (see Priced for rescue), with a federal bailout as the backstop if it sours. The efficiency gains are real. Whether they reach you runs through a layer of companies whose survival depends on them not reaching you.

Three things that cannot all be true

The comforting story needs all three at once: the build-out gets paid off, the whole market gets served, and AI stays cheap. Only two can hold.

  • Pay it off and serve everyone, and prices have to rise, which sends the everyday work to local hardware.
  • Keep it cheap and serve everyone, and the bet is never recouped, and the loss lands on bondholders, pension funds, and, if it concentrates hard enough, the public balance sheet.
  • Keep it cheap and pay it off, and you abandon most of the market, which is the quiet admission that the numbers behind the build-out were never reachable.

Falling token prices are not evidence the story works. They are how the squeeze is applied.