The Price of Intelligence: AI’s Quality-Adjusted Price Collapse

Two 2026 measurements converge on one stylized fact: the cost of a fixed level of AI capability is falling faster than any other measured technology price decline in history — but the measured speed depends almost entirely on the yardstick, and the choice of yardstick is a live methodological fight.

The stylized fact (Epoch AI)

Emberson and Roodman (Epoch AI, September 2026) estimate that across five benchmarks covering mathematics, hard sciences, and games of skill over 2023–2026, the cheapest cost to attain a given benchmark performance fell about 47% per quarter — a 13-fold drop per year. Math problems fall fastest (50–52%/quarter, 16–19×/yr), game-based puzzles slowest (39–43%/quarter). The decline is steepest right after a performance level debuts as state of the art: 66%/quarter (75×/yr) at SOTA, halving to 32%/quarter (4.7×/yr) two years later — consistent with a brief capability premium before competition and efficiency catch up. Canonical datapoint: OpenAI’s o3 reached 75% on GPQA Diamond at roughly 0.0004/question in mid-2026 — a 725-fold fall in under 18 months.1

The methodological novelty is measuring price-to-achieve-performance rather than price-per-token. Reasoning models (post-o1) broke per-token comparisons: they burn far more tokens to extract performance from a base model that is cheaper per token. Epoch borrows the CAISI budget-simulation procedure — take a high-budget benchmark transcript, truncate it at hypothetical per-question token budgets, score unfinished questions at chance — to map each model’s full cost–performance continuum, then takes the Pareto frontier across models at each point in time.2

Prior estimates line up by vintage and method: Appenzeller’s 2024 per-token analysis found ~10×/yr in the three years after GPT-3; Epoch’s own March 2025 milestone analysis found 9–900×/yr across six benchmarks; Ihle’s September 2025 coding-suite analysis found 64–380×/yr; Gundlach et al. (March 2026), the closest methodological sibling, found 5–10×/yr.3

The measurement-careful version (Zhu)

Louis Zhu (Oxford, August 2026) built the statistical-agency-grade version: a monthly panel of 21,024 posted-price observations across 3,208 models and 86 providers, joined to 4,605 benchmark scores through a latent capability index, with a pre-registered validity audit. The paper’s headline is that “how fast is AI getting cheaper” has three different answers:4

  • Matched-model methods (what statistical agencies apply to software): prices fell 0.10 log points/year.
  • Quality-adjusted index: 0.73 log points/year — meaning 87% of the decline is invisible to standard methods, because quality arrives as new models at old prices, exactly the margin matched-model methods cannot see.
  • Per completed task: the buyer’s price stopped falling — reasoning models raised token consumption faster than token prices fell, so the seller’s price (per token) and the buyer’s price (per task) have diverged.

The audit carries the sharpest finding: excluding contamination-flagged benchmarks leaves the capability ranking intact (correlation 0.998) yet moves the price index by 0.49 log points/year. The leaderboard-stability arguments standard in AI evaluation — rankings are robust — therefore offer no defense of economic statistics built on benchmark spacing; an economic statistic can be materially corrupted by validity failures that leaderboard logic waves through.5

The demand side: buyers upgrade rather than economize

Firms are not pocketing the decline as savings. The average price firms actually pay per token stayed roughly stable through the collapse, while the token-weighted median intelligence of the models in use rose from 0.30 to 0.44 over 2025 — price declines are being spent on better intelligence rather than cheaper copies of the same intelligence, the personal-computer/smartphone pattern, except that here both prices and capabilities moved sharply in the buyer’s favor at once.6

Caveats (Epoch’s own)

Benchmark performance is not useful work, and AI companies may be training to the benchmarks (“benchmaxxing”); the frontier measure assumes a relentless cost-minimizing model-switcher, which real users are not; the window is barely three years. And falling unit prices do not imply falling total spend: if a valuable job demands a million benchmark-runs’ worth of cognition, even a penny per run adds up — while the price of new capability (proving hard theorems, as against passing a first-grade math test) is a different curve entirely.7

Open questions

  • Token-unit vs task-unit measurement currently decides the sign of measured real growth in the sector (Zhu) — which unit belongs in productivity statistics?
  • If half a log point a year turns on contamination flags, agency adoption of benchmark-based deflators (the Makridis–Brynjolfsson recommendation) imports the evaluation field’s integrity problems into national statistics. Who audits the auditors?
  • Does the SOTA-premium decay (66% → 32%/quarter) mean the true sticky price is the price of new thought — the frontier itself?

Cross-domain connections

  • llm-inference-provider-landscape — the market structure these prices move through (90 providers, routers vs. hosters, open-weights ~90% below closed); the two pages share the Demirer et al. dataset.
  • ai-economic-growth-complementarity — growth models of AI consume an intelligence price series; this is that series, with Zhu’s warning that the choice of unit can flip its sign.
  • post-scarcity-economics — Jevons’ paradox in live form: collapsing unit price, stable spend, rising consumption as buyers upgrade.
  • eval-awareness-and-grader-orientation — benchmaxxing and contamination are the market-level echo of models optimizing the grader; here they corrupt price measurement rather than capability judgment.

Sources

Footnotes

  1. Luke Emberson and David Roodman 2026 — The plunging price of thought ↩

  2. Luke Emberson and David Roodman 2026 — The plunging price of thought ↩

  3. Luke Emberson and David Roodman 2026 — The plunging price of thought ↩

  4. Louis Yiven Zhu 2026 — The Price of Intelligence: A Quality-Adjusted Price Index for AI Services ↩

  5. Louis Yiven Zhu 2026 — The Price of Intelligence: A Quality-Adjusted Price Index for AI Services ↩

  6. 2025 — Demirer Etal Emerging Market Intelligence ↩

  7. Luke Emberson and David Roodman 2026 — The plunging price of thought ↩