All prices in this post are nominal USD list prices per million tokens, as published by the providers or collected by third-party trackers. "Comparable capability" claims are approximate — benchmarks disagree, and a 2025 model that matches GPT-4 on one task may not match it on another. Figures for 2026 model pricing come from aggregator sites and should be treated as indicative.
The Headline Fact
In March 2023, OpenAI launched GPT-4 at $30 per million input tokens and $60 per million output tokens. Using the common 80/20 input-output blend, that is roughly $36 per million tokens — the going rate for state-of-the-art machine intelligence at the time.
By early 2025, Google’s Gemini 2.0 Flash sold for $0.10 input / $0.40 output — and on most public benchmarks it matched or beat the original GPT-4. Blended, that is a price decline of roughly 200x in under two years for a fixed level of capability. Even if you distrust benchmark equivalence and haircut the comparison aggressively, the decline is still two orders of magnitude.
This is not one company’s pricing strategy. Epoch AI tracked the price of reaching fixed performance thresholds across many benchmarks and found declines of 9x to 900x per year depending on the task, with a median around 50x per year — and the pace accelerated after January 2024. For the specific milestone of matching GPT-4 on PhD-level science questions, prices fell about 40x per year.
No mainstream input has ever gotten cheaper this fast. Transistors under Moore’s law doubled per dollar every two years or so — about 1.4x per year. Solar panels fell about 20% per doubling of capacity. LLM inference at fixed capability is falling 10x to 100x per year. If you are building a business on top of this input, the price you pay today is trivia; the trajectory is the whole game.
Why Prices Collapse
The collapse is over-determined — several independent forces each contribute a multiple, and they compound. Here are the main drivers, with rough magnitudes. The magnitudes are approximate and overlapping; they do not multiply cleanly.
| Driver | Type | Rough contribution | What it does |
|---|---|---|---|
| New GPU generations | Hardware | ~2–4x per generation | Hopper → Blackwell-class parts deliver several times more inference throughput per dollar, especially at low precision (FP8/FP4). |
| Quantization | Algorithms | ~2–4x | Serving weights at 8-bit or 4-bit instead of 16-bit cuts memory and compute per token with small accuracy loss. |
| Distillation & smaller models | Algorithms | ~5–10x | Training compact models on the outputs and data recipes of larger ones. A 2025 "flash"-class model matches a 2023 giant at a fraction of the parameters. |
| Mixture-of-experts sparsity | Algorithms | ~3–10x | Only a small fraction of parameters activate per token. DeepSeek V3 has 671B total parameters but activates ~37B per token. |
| Serving-stack engineering | Algorithms | ~2–5x | Continuous batching, paged KV caches, prompt caching, and speculative decoding raise GPU utilization from "embarrassing" to "respectable". |
| Competition | Market | Compresses margins to the floor | Open-weight models (Llama, DeepSeek, Qwen) let anyone with GPUs serve near-frontier capability. Undifferentiated tokens get priced near marginal cost. |
A few of these deserve a sentence more.
Hardware compounds slowly; algorithms compound fast. GPU price-performance improves at Moore’s-law-like rates — meaningful, but it explains maybe one of the ten-fold declines per two years. The rest is software: model architectures that need fewer FLOPs per unit of capability, and serving stacks that waste fewer of the FLOPs they have.
Sparsity changed the pricing equation structurally. A dense model pays for every parameter on every token. A mixture-of-experts model stores a large brain but consults only a sliver of it each step. This is why DeepSeek could serve a GPT-4-class model at roughly $0.27/$1.10 per million tokens in late 2024 — around a hundredth of GPT-4’s launch price — and, by its own (disputed but directionally plausible) accounting, still run inference at a profit.
Competition converts cost declines into price declines. Falling cost alone does not force a price cut; competition does. Every provider knows that a customer’s switching cost between “OpenAI-compatible API endpoints” is an afternoon of engineering. When an open-weight model reaches a capability tier, the price of that tier collapses to near the marginal cost of serving it, because dozens of hosting providers race each other down.
The Chart: Watching a Price Collapse
The chart below plots the blended (80% input / 20% output) list price for models at roughly GPT-4 capability, from GPT-4’s launch to early 2025. The Y-axis is logarithmic — each gridline is a fixed multiple, so a straight line means a constant rate of collapse.
Two things stand out. First, the line is close to straight on a log scale: the decline is exponential, not a one-off correction. Second, the steepest segment comes from entrants, not incumbents — the DeepSeek and Gemini Flash points are where the curve falls off a cliff. Incumbents cut prices along a schedule; competitors reset the schedule.
Where does it stand in 2026? Aggregator data (indicative, not official) puts the current frontier tier at $5/$25 (Claude Opus class) down to $1.75/$14 (GPT-5.2), mid-tier at around $3/$15, and economy tiers at $0.50/$3.00 (Gemini 3 Flash) down to $0.14/$0.28 (DeepSeek V4 Flash). The umbrella holds at the top; the floor keeps dropping.
Price Is Not Cost: The Margin Debate
A price chart tells you what buyers pay. It does not tell you what serving costs — and that gap is one of the most consequential unknowns in the AI economy.
What does a serving node actually cost? Rough public numbers: an 8x H100 server costs roughly $200k–$400k to buy, or $2–$7 per GPU-hour to rent depending on provider and commitment; each H100 draws up to 700W under load, and cooling, facilities, and maintenance add materially on top. Against that cost, throughput is the whole story. A well-batched node serving a mid-size model can emit on the order of tens of thousands of tokens per second across concurrent requests; a badly utilized one, a hundredth of that.
That makes utilization the hidden variable in every margin claim. The same hardware serving the same model can imply a cost per million tokens that differs by 50x between a full batch at peak and an idle node at 3 a.m. This is why self-hosting analyses find you need thousands of conversations per day before running your own GPUs beats API prices: the APIs are pooling everyone’s demand to keep batches full.
So are the labs selling below cost? Both narratives circulate, and honest observers admit the public evidence cannot settle it:
- The bull case (prices are above marginal cost): batched inference on modern hardware is genuinely cheap; DeepSeek published (contested) figures implying healthy margins at prices far below Western labs’; and the relentless price cuts look like cost pass-through, not deepening subsidy — you do not repeatedly cut prices on a product you lose money on unless costs fell first.
- The bear case (prices are subsidized): the labs are in a land-grab for developers and data, funded by tens of billions in venture and hyperscaler capital; free consumer tiers and flat-rate subscriptions are almost certainly served below cost for heavy users; and marginal cost excludes the training runs and datacenter buildouts, which someone eventually has to recoup.
The likely resolution is boring: paid API tokens are probably gross-margin positive at the margin, while the businesses as a whole burn cash on training, free tiers, and capacity built ahead of demand. But nobody outside the labs’ finance teams knows, and every margin figure you read — including the ones above — is an estimate built on estimates.
Interactive: The Inference Cost Time Machine
Numbers per million tokens are abstract. What matters is your workload’s monthly bill — and how violently that bill has changed across three years of pricing. Set up a workload below, pick a 2026 price tier, and see what the identical workload would have cost at 2023 and 2024 prices.
Inference cost calculator
Monthly cost = requests/day × 30.4 × (input tokens × input price + output tokens × output price). List prices, no caching or volume discounts.
This workload is — cheaper at 2026 economy-tier prices than at GPT-4's 2023 launch prices.
Try the default workload — 5,000 requests a day, a 2,000-token prompt, a 500-token answer. At GPT-4’s 2023 prices it was a five-figure monthly line item that needed a director’s sign-off. At 2026 economy-tier prices it costs less than a team lunch. Entire categories of product — summarize every support ticket, screen every resume, caption every image — crossed from “not economically viable” to “rounding error” without the products changing at all. The price moved; the use cases followed.
Jevons’ Paradox: Cheaper Tokens, Bigger Bills
In 1865, the economist William Stanley Jevons observed that more efficient steam engines did not reduce Britain’s coal consumption — they exploded it, because efficiency made coal-powered everything economically viable. Token prices are the coal of this cycle.
The aggregate numbers are stark. Per-token prices fell more than 90% from 2023 levels, yet enterprise spending on generative AI grew from roughly $1.7B in 2023 to $11.5B in 2024 to $37B in 2025 — a 22x increase in two years, during the steepest price collapse in the technology’s history. Falling unit prices did not shrink the market. They were the mechanism by which the market grew.
Agents turbocharge the effect. A chat interaction is one prompt, one reply. An agent loops: it plans, retrieves, calls tools, reads the results, self-checks, retries, and escalates — dozens or hundreds of model calls per task, each carrying accumulated context. Microsoft Research measured agentic workloads consuming on the order of 1,000x more tokens than a chat interaction for an equivalent task. Cheaper tokens are precisely what makes those architectures affordable to attempt, so each 10x price decline unlocks workloads that consume 100x the tokens.
For any individual company, the budget experience is therefore paradoxical:
- Per-unit cost falls — the finance team expects the AI line item to shrink.
- Viable use cases multiply — every team finds a workload that now pencils out.
- Architectures get hungrier — a wrapper becomes a RAG pipeline becomes a multi-agent system.
- The bill goes up — by 2026 the FinOps community had a name for the collective realization: the “Great Token Panic.”
What It Means
For app builders: the moat question
If your product is a thin wrapper — a prompt, a model call, a UI — then your cost of goods is collapsing, which sounds great until you notice your competitors’ costs collapse identically, and your customers’ cost of self-serving collapses too. Thin margins on top of a deflating commodity are not a business; they are a countdown.
The defensible positions sit where cheap tokens are an input, not the product: proprietary data and feedback loops the model consumes, workflow integration deep enough that switching hurts, distribution, and accountability (compliance, evaluation, uptime guarantees) that enterprises pay for regardless of the token price. Cheaper inference makes those moats more valuable, because it lets you spend lavishly — verification passes, multi-model ensembles, agentic retries — on quality your wrapper competitors cannot match at the same retail price.
For the capex debate: deflation as the bull and the bear case
Hyperscalers and labs are committing hundreds of billions of dollars to AI datacenters. Token deflation cuts both ways through that math. The bear reading: how do you earn a return on capacity whose output price falls 10x a year? Revenue per GPU-hour is a melting ice cube. The bull reading: Jevons. Demand has so far grown faster than prices fell — total spending rose 22x while unit prices dropped 90% — so aggregate inference revenue can rise even as each token gets cheaper, and the cheapest capacity keeps winning the marginal workload.
Which reading wins is the open question of the AI investment cycle, and it likely resolves differently for different players: owners of scarce power, land, and interconnects capture deflation-resistant rent; owners of rapidly depreciating mid-tier GPUs may not. What is already clear is that betting on token prices staying high is the one position with no historical support.
For users: intelligence too cheap to meter, almost
For individuals, the practical consequence is that yesterday’s frontier keeps arriving in free tiers and $20 subscriptions. The capability you paid API rates for in 2023 is bundled into your email client by 2026. The scarce resources shift from access to intelligence toward the things intelligence cannot substitute for: your data, your attention, your judgment about when the model is wrong.
Key Takeaways
References
- OpenAI. "API Pricing." GPT-4 launch pricing ($30/$60 per 1M tokens, March 2023) and subsequent GPT-4 Turbo / GPT-4o pricing. openai.com.
- Ng, Andrew. Post on X comparing GPT-4 launch pricing ($36/1M blended, March 2023) with GPT-4o pricing ($4/1M blended, August 2024). x.com.
- Epoch AI. "LLM inference prices have fallen rapidly but unequally across tasks." Price declines of 9x–900x/year across benchmarks; 40x/year for GPT-4-level performance on PhD-level science questions. epoch.ai. Retrieved July 2026.
- DeepSeek. "Models & Pricing." DeepSeek API documentation. api-docs.deepseek.com.
- Google. "Gemini Developer API Pricing." ai.google.dev.
- Introl. "Inference Unit Economics: The True Cost Per Million Tokens." GPU node costs, rental rates, utilization break-evens. introl.com. Retrieved July 2026.
- Fortune. "Tokens are getting cheaper, but companies are spending even more on AI as a result." June 2026. fortune.com.
- IntuitionLabs. "AI API Pricing Comparison (2026)" — 2026 tier pricing for GPT-5.x, Claude 4.x, Gemini 3.x (indicative aggregator data). intuitionlabs.ai. Retrieved July 2026.
- Jevons, W. Stanley. The Coal Question. Macmillan, 1865.
Michael Wan Interactive Insights