ManchesterStory Group LLC

08/05/2026 | Press release | Archived content

Open-Weight Models Are Driving Down the Cost of AI

The last few weeks have brought a lot of major developments in AI. Anthropic's Claude Opus 5, released on July 24, now narrowly tops the independent Artificial Analysis Intelligence Index with a score of 61, ahead of its own Fable 5 at 60 and OpenAI's GPT-5.6 Sol at 59. Notably, Opus 5 reaches Fable-level intelligence at roughly half the blended token price. More interesting to us is how quickly the open-weight models are closing the performance gap, and how much cheaper they are.

ManchesterStory continues to pay close attention to the rapid development of AI, as it impacts the whole startup ecosystem.

The last few weeks have brought a lot of major developments in AI. Anthropic's Claude Opus 5, released on July 24, now narrowly tops the independent Artificial Analysis Intelligence Index with a score of 61, ahead of its own Fable 5 at 60 and OpenAI's GPT-5.6 Sol at 59. Notably, Opus 5 reaches Fable-level intelligence at roughly half the blended token price.

More interesting to us is how quickly the open-weight models are closing the performance gap, and how much cheaper they are.

An open-weight model is one where the developer publishes the trained parameters, or weights, so anyone can download the model and run it on their own hardware rather than only renting access through a paid API. However, the largest open-weight models require enormous compute to self-host, so in practice most teams still consume them through a hosted provider, simply at a materially lower price than proprietary models.

Moonshot's Kimi K3, whose weights were released for public download on July 27, scores 57 on the same index, within three points of Fable 5, at a blended cost of $2.31 per 1M tokens - roughly a third of Fable's $7.70 and well below Opus 5 at $3.85. Z.AI's GLM-5.2 shows how much further down the curve goes, scoring 51 at just $0.90 per 1M tokens.

Open-weight models are narrowing the intelligence gap while sharply undercutting proprietary pricing. Kimi K3 scores within two points of GPT-5.6 Sol at nearly half the blended cost, while GLM-5.2 delivers a score of 51 for just $0.90 per 1 million tokens.

From Token-Maxxing to Thrift-Maxxing

The price of tokens has become the industry's central topic. Companies spent a year competing to consume as many tokens as possible and now shop for the cheapest model that can finish the job. The Wall Street Journal has given the behavior a name: thrift-maxxing. As Cursor's field CTO Mike Saeks put it, running every task on a frontier model is "like driving a Lamborghini to go to the grocery store to pick up milk."

The savings are not marginal. Cursor priced out building a web browser from scratch: a little over $10,000 to run the whole job on GPT-5.5, versus $1,339 splitting it between Cursor's own Composer model and Anthropic's Opus 4.8.

This is not a story about abandoning US labs. It is a story about routing each task to the cheapest model that can handle it and reserving the expensive models for the work that justifies the price. Cursor built its own Composer 2 coding model on Kimi foundations, and the legal AI startup Harvey fine-tuned GLM-5.2 in-house before giving it a button to call Fable 5 when a task turns out to be genuinely hard. Coinbase said in June it had halved its AI spending by pushing staff toward Kimi and GLM. DoorDash routes what its CTO calls "lower-level work" to Kimi and Airbnb has used Alibaba's Qwen for customer service.

Why we see this as positive for the ecosystem

We view the cost compression around near-frontier models as a positive for the startup ecosystem. Capability close to the frontier at a lower price means early-stage companies can build and ship considerably more on the same amount of capital, and also improves the margin profile. We have seen this in our own portfolio, where companies have moved workloads to open-weight models and meaningfully reduced their AI spend without a drop in output quality. More established companies like Airbnb and NVIDIA are following the same path of often running a mix of frontier and open-weight models.

Conclusion

The direction of travel is hard to miss. Intelligence is getting cheaper faster than most people forecast, the gap between the best model and a good-enough model is narrowing, and the cost of switching between them keeps falling. For the startup ecosystem, it is hard to see much of a downside. We will keep watching where the price curve goes from here, and in particular whether the open-weight models can hold this pace at the frontier. Our sense is that we are closer to the beginning of this trend than the end of it.

Sources

Artificial Analysis Intelligence Index v4.1 (July 2026); The Next Web, "A million words costs $50 from Anthropic and 87 cents from DeepSeek. Corporate America noticed." (July 27, 2026); The Wall Street Journal; Fortune; IDC; Bloomberg Intelligence.

ManchesterStory Group LLC published this content on August 05, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on August 18, 2026 at 19:57 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]