(20) AI Buyers No Longer Pay for Capability They Cannot Use
The market has found a ceiling
AI buyers no longer pay a frontier premium for capability that most of their workload cannot use. Claude Fable 5 is the clearest early test. Ramp’s August 2026 AI Index found that, during the preceding month, Fable accounted for 6% of the Anthropic tokens purchased through Ramp’s token-management product and 11.4% of model-attributed Anthropic spending.
Those figures are narrow. Ramp says the sample skews toward technology companies, and one month is too little to declare a commercial failure. Fable was unavailable to all users from June 12 until July 1, while Opus 5 arrived on July 24, so the early data cannot isolate price from availability or internal substitution. The gap between token share and spending share also shows that some customers will pay the premium. Even so, Fable has not become their default model.
Anthropic charges $10 per million input tokens and $50 per million output tokens for Fable. Opus 5 costs half as much, while Sonnet 5 is cheaper again. Anthropic’s own model guide tells undecided customers to begin with Opus 5 for complex agentic and enterprise work, reserving Fable for workloads that need the highest available capability. The product catalogue now encodes the buyer’s hesitation.
Good enough has become cheap
The market around Fable matters more than its benchmark position. Mert Demirer, Andrey Fradkin and Nadav Tadelis studied business-to-business inference data from OpenRouter for the Journal of Economic Perspectives. They found that the price of a given level of intelligence had fallen roughly a thousandfold, with open-source models costing about 90% less than comparable closed models. No single model dominated every use case.
The longer trend points in the same direction. Stanford’s 2025 AI Index calculated that the cost of querying a model at GPT-3.5-level MMLU performance fell from $20 per million tokens in November 2022 to $0.07 by October 2024. A capability that once required a premium model became a low-cost input in less than two years.
The “90% of tasks” claim should be treated as a purchasing heuristic, not a universal statistic. Every company has a different workload. Once several models clear the required quality threshold for a routine task, extra intelligence produces little additional value. Price, latency and predictability take over.
Routing shrinks the market for the best model
Production data already shows this division of labour. Vercel reported that open-weight models processed 29% of tokens on its AI Gateway in June 2026 while accounting for less than 4% of spending. High-volume work moved toward cheaper models, while frontier providers retained most spending and dominated consequential uses such as coding and back-office agents.
This is the non-obvious problem for Anthropic. Better routing can make Fable more valuable while reducing its share of tokens. A cheaper model handles repeatable work; Fable receives the uncertain, long-running problems where a mistake or failed attempt costs more than the inference bill. Its role becomes closer to an escalation layer than a default utility.
Opus 5 makes that segmentation sharper. It arrived at half Fable’s token price and Anthropic recommends it as the starting point for most agent workloads. The company may therefore be limiting Fable’s volume with a product that protects the wider Claude franchise. Low Fable usage can coexist with rising Anthropic adoption, which Ramp also recorded in July.
The expensive model still has a case
Per-token prices can mislead. A stronger model may finish in fewer turns, avoid a retry or reduce human review. On Anthropic’s DeepResearch Bench II, Fable at low effort was more accurate and about 10% cheaper per completed task than Sonnet 5. Yet on Anthropic’s SWE-bench Pro subset, Opus 5 matched Fable within run-to-run noise at about 60% of its cost. Anthropic’s evidence supports both sides because the ranking changes with the workload.
That is why Fable’s adoption challenge cannot be solved by another leaderboard win. Anthropic must show where its extra capability lowers the cost of an accepted result, including retries, review time and the consequences of failure. Buyers should make the same calculation on their own traffic and send only the hard tail to Fable. The premium model earns broad revenue when its advantage is measurable; it does not need broad usage.
Sources
-
Ara Kharazian, “Cracks in the AI Thesis” https://ramp.com/data/ai-index-august-2026
-
Anthropic, “Models overview” https://platform.claude.com/docs/en/about-claude/models/overview
-
Anthropic, “Optimizing for cost and intelligence” https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence
-
Anthropic, “Redeploying Fable 5” https://www.anthropic.com/news/redeploying-fable-5
-
Eric Dodds and Jerilyn Zheng, “Open-weight models surge to 29% of volume, price per token flattens” https://vercel.com/blog/ai-gateway-production-index-july-2026
-
Mert Demirer, Andrey Fradkin and Nadav Tadelis, “The Emerging Market for Intelligence: How Firms Buy and Sell AI” https://www.aeaweb.org/articles?id=10.1257/jep.20261506
-
Stanford Institute for Human-Centered Artificial Intelligence, “Artificial Intelligence Index Report 2025: Chapter 1 — Research and Development” https://hai.stanford.edu/assets/files/hai_ai-index-report-2025_chapter1_final.pdf