When Chinese company Moonshot AI revealed that its Kimi K3 model performed just as well, if not better, than the newest OpenAI and Anthropic models, it added fuel to the fire on the debate between open-weight and closed, proprietary models.
See Also: OnDemand | Security Operations in the Age of AI
The Trump administration is openly weighing whether to ban Chinese models, and American frontier labs like OpenAI and Anthropic are calling for a slowdown in artificial intelligence development.
There is no doubt that Chinese open-weight models are becoming better. Based on Moonshot AI's internal benchmarking, Kimi K3 came in a very close second to OpenAI's GPT-5.6 Sol in the Terminal Bench 2.1 coding test and was third behind GPT-5.6 Sol and Anthropic's Fable 5 in the DeepSWE benchmark. Other recent releases from DeepSeek and Alibaba also showed high scores in the same benchmark tests.
Enterprises have begun seriously looking at using Chinese open-weight models, which many independent evaluators estimate are between four and eight months behind closed frontier models. Most observers note that while American frontier model firms make a compelling capability argument, the driving force of adoption is cost.
In the AI world, being "open" means many different things. Many software engineers talk about open-source projects, which require researchers and developers to release as much information as possible about the model, allow other labs to download and enhance the codebase. Most AI models, even those from Chinese companies, are open weight, which offers outside teams trained parameters or weights so these can be downloaded, run on their own servers and fine-tuned. The training data, code and methodology remain locked.
While no benchmark tracks the timeline of Chinese models - these are private companies after all - some researchers have developed some ways to estimate how far behind foreign open-weight models are compared to the leading frontier labs like OpenAI and Anthropic. The consensus is that Chinese open-weight models are around half a year behind, but there's no real indication that this gap is closing.
Chris Canal, founder of AI evaluation company Equistamp, told ISMG in an interview that he puts Chinese models "around six to eight months in terms of general capabilities," based on how far behind a Chinese model announces its benchmark score compared to a leading model.
"Although they are still behind, they're not that far behind," Canal explained.
Alexander Barry, senior researcher at the non-profit AI research firm Epoch AI, said the gap between open-weight and closed-source models tends to fluctuate. Barry pointed out that there hasn't been a similar "DeepSeek moment" in the past few model releases, referring to the market's reaction when DeepSeek's LLMs showed an ability to beat OpenAI's models. Now, it's almost expected that Chinese and Western models trade places on AI performance leaderboards.
If you ask Anthropic or OpenAI why Chinese models are becoming very performant, very fast, they will most likely point to instances of "illicit distillation." Anthropic and OpenAI have separately accused Moonshot AI, DeepSeek and Alibaba of using unlawfully closed models to train their new models. This allows anyone using this method to basically use the same training data as proprietary LLMs, by asking the model questions and then learning from its answers, without needing to download the same datasets.
But there is one area where Chinese open-weight models consistently beat American models: cost.
Performance at Half the Price
Most organizations implementing AI projects have had to contend with the rising cost of using it. As models become larger and reason more, they require more tokens to answer questions, complete tasks, or remember context.
When DeepSeek arrived on the scene, people were shocked that it could do the same things as OpenAI's o1 model, but at a fraction of the price.
Anthropic currently prices its top two models, Claude Fable 5 and Mythos 5, at $10 per one million input tokens and $50 per one million output tokens. The newest flagship OpenAI model, GPT-5.6 Sol, costs $5 per one million input tokens for short prompts and $10 per one million input tokens for longer contexts.
On the other hand, Kimi K3 costs between $0.30 and $3 per one million input tokens and $15 per one million output tokens. Alibaba's newly released Qwen 3.8 Max arrives at $2 for inputs and $6 for outputs, according to OpenRouter.
John Larson, president and chief AI officer at Babel Street, told ISMG that market forces are going to begin putting pressure on closed models.
"I think that markets are going to start to dictate investment around these," Larson said. "Foundational AGI models are extremely expensive, very computationally expensive and there's going to be a market need that needs to be filled, which is why I think Chinese models have risen."
This may already be having an effect. OpenAI announced in July that it slashed prices for two smaller models, GPT 5.6 Luna by 80% and Terra by 20%.
Almost all of the industry observers ISMG spoke to conceded that the adoption of Chinese open-weight models is largely driven by affordability. Though organizations understand that there may be some security or bias risks within these models, in an increasingly multi-model environment where enterprises want access to the best models for specific use cases, the fact that a Kimi K3 performs similarly to a model almost twice its price makes it very compelling to use.