07/30/2026 | Press release | Distributed by Public on 07/30/2026 12:13
Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus across the industry. How do we get the best possible performance per token invested, and the best real customer outcome per dollar invested?
To build a frontier firm, you have to optimize frontier performance against cost. Choosing where you want to sit on that curve is critical. By co-optimizing your models, harnesses, and RLEs you can pick a point on the curve that suits your firm.
In most cases, frontier generalist models aren't necessary for every task. By tuning models for a specific product, you can maintain or even exceed frontier performance, while reducing token costs dramatically.
This is where we have focused our MAI hill-climbing machine over the last quarter, and the results are pretty cool. This week we released MAI-Cyber-1-Flash optimized for our MDASH harness.
Together, the system landed at No.1on the leading CyberGym benchmark - beating Mythos by 12ppts - at 50% of the cost. And remarkably, we serve it on H100s too.
It was designed to handle up to 90% of tasks efficiently, so that MDASH can reserve the largest and most expensive models in our fleet (in this case GPT 5.4) for the 10% of exceptionally hard problems that truly need them.
As Satya mentioned today in our Q4 Earnings call, since last quarter, we've shipped more than a dozen new models across image, voice, transcription, coding and security, and they're already powering many of Microsoft's most widely used products to maintain or improve quality while using significantly fewer tokens, in many cases saving 50-90% of GPU costs:
And what's more, by co-designing our models with our own silicon, we are seeing 40% better performance per watt running MAI models on Maia 200.
But the benefit is not only cost. It's resilience. Every business now must assume that any one model it depends on could disappear, through a security incident, a business or policy misalignment, or a geopolitical shift.
Every model in a product or agentic system should be substitutable, and that's only possible when you build the harness, context, memory and action space independently of a single model family. That's the hill-climbing machine we've built.
We think this is the beginning of a genuinely new performance curve. Its shape represents a system rather than a model, and traversing this curve delivers better quality, lower cost, and more choice.
This has been a summer of hard but wonderful work by the team. We are keenly aware of how early this is, and of how much we still have to learn. But the direction is clear, we are hill-climbing to move the frontier on the cost-to-outcome curve, and we will keep sharing what we learn along the way. There is much more to come.