// TOM'S HARDWARE US — HARDWARE & GADGET
Frontier AI faces pricing reckoning as token volume explodes 25-fold — mid-tier models deliver 90% of flagship capability at one-sixth the cost
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
AI development might not be the wild west it was when ChatGPT burst onto the scene a few years ago, but it's still very much a frontier, with no clear boundaries and few yardsticks. But for AI developers on the frontier, they're pulling hard towards dual goals of ever greater intelligence and ever cheaper per-token pricing, and it's leading to a real back-and-forth of who's truly ahead, with some winners only holding the top spot for a few hours.
Although Anthropic's Claude Fable and Opus models have been consistently competitive at the very top of the intelligence charts, they're also some of the most costly to use. For more general use, some are paying closer attention to the "Pareto Frontier," where peak intelligence and minimal cost reach the pinnacle, and there the competition is fierce and ever-changing.
Hot off the screeching reversal of companies' tokenmaxing plans earlier this year, this increased focus on getting the cost of AI down has left us running headfirst into Jevons paradox again, too. As token costs for high-intelligence models have come down, token usage has exploded over 25 times in the past year, and doubled in the past month alone.
People may not want to spend more on AI, but they appear to be using a lot more of it when they can afford to.
Despite its radical and rapid ascension, the big winners in the AI industry haven't changed much since its inception. It may have had a few penny drop, "Deepseek moments," where there's been a frenzied scramble by everyone to get ahead of some new threat, but by and large OpenAI and Anthropic have been scuffling at the top of the intelligence pile, Google and Meta have been bouncing around the more efficient and cost-effective middle, and xAI's Grok has been there in the background, grabbing headlines for all the wrong reasons.
That's largely still the state of play in September 2026. Although benchmarks are gamed during model design and real-world use is more representative of actual real-world use, Anthropic's best are still considered by most to be the smartest. Fable 5.1, Fable 5, and Claude Opus all rank in the top four of ArtificialAnalysis' Intelligence Index test, as does OpenRouters and BenchLM even have them take all the podium spots.
While ahead, though, Anthropic's models don't hold an enormous lead. Fable 5.1 might score a 66 on ArtificialAnalysis' benchmark, but OpenAI's GPT 5.6 Sol (max) manages a 61. Grok 4.6 (high) and Kimi K3 (max) are capable of scores above 60, and the new Meta Muse Spark 1.3 (max) can hit 62 - though we don't have cost comparison pricing for it yet.
The same is true across other benchmarks from other companies.
But where the top models nudge each other back and forth with light tweaks and slight bumps in capability, there's much greater distinction in the mid-range. And not on intelligence, but on price.