Token prices are falling. Your bill isn't.
Two facts coexist, and they seem incompatible.
Token prices are collapsing. And software publishers’ AI budgets are growing.
The contradiction is only apparent. It comes from confusing two distinct quantities: the unit price your provider charges, and the number of units your product consumes.
The first is falling. The second is exploding.
Understanding this gap changes how a publisher builds their P&L.
The deflation is real, and dramatic
Let’s start by validating the first fact.
Epoch AI, a research institute specializing in AI trajectories, measured how fast the price needed to reach a given performance level falls. The method is published, the code is reproducible. Six benchmarks, three years of data.
The result is clear-cut. To reach GPT-4’s level on PhD-level science questions, the price has been divided by 40 every year. Depending on the performance tier targeted, the drop ranges from 9x to 900x a year, with a median of 50x.
This deflation has even accelerated. The fastest drops came after January 2024. Restricting the analysis to that period, the median jumps from 50x to 200x a year.
No computing commodity has ever fallen this fast. Not compute during the microprocessor revolution, not bandwidth during the dot-com boom.
The first fact is established. On to the second.
The tokens you don’t see
A methodological detail from Epoch AI deserves attention, because it contains the whole explanation.
The researchers excluded reasoning models from their token-price analysis. Their reason is straightforward: these models generate far more tokens than others, which makes any unit-price comparison at equal performance misleading.
Take a moment with that. The reference institute on the topic removes an entire category from its calculation, because price per token no longer means much there.
The mechanism is simple. A reasoning model produces two kinds of tokens. The ones you read in the answer. And the ones it generates to think, which you don’t see but do pay for. These thinking tokens vary by an order of magnitude from one model to another, on an identical request.
Add agentic architecture on top. An agent that breaks down a task, calls tools, evaluates its results, and retries consumes far more than a simple completion. Your unit cost falls. Your unit volume doesn’t follow the same curve.
When the cheapest model costs more
This reasoning has just been quantified.
A team from Stanford, UC Berkeley, CMU, and Microsoft Research evaluated eight leading reasoning models on nine tasks — competition math, science questions, code generation, multi-domain reasoning. Goal: compare sticker price to actual cost incurred.
The result breaks the intuition. In 21.8% of pairwise model comparisons, the one with the lower list price actually costs more in practice. The gap reaches 28x in extreme cases.
A concrete example from the study: Gemini 3 Flash lists a rate 78% lower than GPT-5.2. Its real cost, measured across all tasks, comes out 22% higher.
The authors go further. They consider predicting a request’s real cost from its price and content to be an open, unsolved problem.
That’s worth remembering. Researchers with the data and the measurement tools conclude that the math isn’t reliable as it stands.
This work is available as a preprint on arXiv. It hasn’t yet been peer-reviewed.
What this means for a software publisher
The problem isn’t that AI is expensive. It’s that the unit you’re billed on isn’t the one you control.
You control requests, users, use cases. You’re billed in tokens. Between the two sits a coefficient that depends on the model chosen, how long it reasons, how many iterations your agents run. You don’t know this coefficient in advance. It changes with every model update.
A second phenomenon compounds this invisible coefficient: scale. Going from a few thousand tasks to a few million isn’t just the same problem multiplied by a thousand. Rare cases become frequent, long or complex requests that stayed marginal on a test panel become the norm, and the tokens-per-task coefficient measured at small scale doesn’t hold.
Three practical consequences:
- A pricing grid pegged to your provider’s list price is wrong. By an unknown factor, potentially an order of magnitude.
- Choosing between models based on sticker price can cost you more. That’s exactly what the Stanford study measures.
- Unit-price deflation isn’t protection. It’s real, but it moves on a variable your total costs don’t follow.
The question to work through isn’t “when will the token be cheap enough.” It’s: which unit do you want to build your business model on, and which one can you actually forecast.

Sources
- Cottier, Snodin, Owen & Adamczewski, LLM inference prices have fallen rapidly but unequally across tasks, Epoch AI, March 2025 — epoch.ai/data-insights/llm-inference-price-trends (CC BY)
- Chen, Zhang, He, Stoica, Zaharia & Zou, The Price Reversal Phenomenon: When Cheaper Reasoning Models End Up Costing More, preprint arXiv, March 2026 — arxiv.org/abs/2603.23971
Mastering your business model means mastering your cost unit. Agora Software provides a hosted multi-agent platform (on-premise or sovereign cloud) with costs known in advance, independent of token prices.
Bring AI into your software with Agora Software.
Let's talk