The race to build the smartest LLM is over. Frontier labs are optimizing on a new metric now.
Arnav Jaitly
Aug 23, 2026
For the last 3 years, every new model headline relied on reporting benchmark scores . Off late, that trend seems to be fading away. I suspect this is because consumers don’t have a material change in their lives due to a 3% change on some obscure benchmark. What impacts them is the price they pay to get more of their work done. So the labs are now finetuning their token-economics more than their hyperparameters.
OpenAI cut GPT-5.6 Luna's price 80% in July. Yesterday it cut Sol's price again, explicitly to undercut Claude Opus 5. And if you’re wondering if Anthropic is silently on this, then you’re wrong. They are fighting the same war through acquisitions. It's reportedly in talks to pay $6B for Decart, a startup built to squeeze more usable output out of the same compute. And that is the new metric. How cheap can you make the smartest models, and how fast can you push that number down again.
Its not just the frontier labs playing this game. AMD made the same bet twice this month. It bought Taalas, a company working to print a model's weights directly into silicon so the chip never has to fetch them from memory. Taalas claims their chips can output 17,000 tokens per second where a standard GPU manages a few hundred. They have also partnered with Cerebras, whose new chip already claims 30x GPU speed, to fold that same advantage into its own hardware roadmap.
Other in the AI supply change want in on this party too. Samsung just raised advanced chip prices up to 15%. IBM is paying Together AI $240M for a cluster built for one stated reason: cheaper inference for companies trying to control AI spend.
The application layer is benefitting the most out of this pricing war. Higgsfield AI has quadrupled its valuation to $5.4B in eight months, backed by $700M in real revenue from 30 million users. It didn't win by having the smartest AI model. It is winning by making AI video actually affordable and usable at scale, and investors are pricing in more of exactly that.
So yes, you don't need the smartest model. You just need the cheapest one that's smart enough.Exciting time to be a token consumer!