A lot happened in AI this week, and if you read the headlines separately it looks like noise. Read them together and there is one story. Chinese labs are giving away frontier-class models for free, and the only thing they cannot give away, or even buy enough of, is compute. That is the thesis. Here is the evidence.
Start with Moonshot AI. This week the Chinese lab released Kimi K3, a 2.8 trillion parameter model with open weights, meaning anyone can download and run it. Fortune and CNBC both called it a new DeepSeek shock, and the benchmark numbers back that up. On coding benchmarks, K3 beat everything except the very top US frontier models. Not everything Chinese, not everything open, everything. A free download is now within reach of the models that cost twenty dollars a month.
Open weights as an export strategy
Then Alibaba matched the move. It announced Qwen3.8 Max, a 2.4 trillion parameter model going open weight on July 21, and claimed it trails only the single best frontier model available. Two trillion-parameter-class releases, both open, in the same week, from two different Chinese companies. That is not a coincidence, it is a pattern.
Think about what open weights do to the US labs' business model. OpenAI and Anthropic sell subscriptions and API access. Their moat is that the best model lives behind their paywall. Every time a Chinese lab releases something 95 percent as good for free, that paywall gets a little harder to justify, especially for businesses running high-volume workloads where the affordability gap is the whole decision. When you cannot win the subscription game, the next best move is to make sure nobody can, and to become the default infrastructure everyone builds on instead. Open weights are China's export strategy the way cheap solar panels were. The markets noticed, semiconductor stocks sold off on the news.
The bottleneck is silicon, not ideas
Here is the twist that proves the thesis. Kimi K3 was so popular that Moonshot paused new subscriptions. Not raised prices, paused signups entirely. The company is compute constrained, and US export controls on Nvidia chips make that worse, not better. Moonshot can give the weights away for free, but it cannot serve the demand for its own product. The ideas travel instantly. The chips do not.
You can see the same bottleneck everywhere else in this week's news. Nvidia's next-generation Kyber rack system is reportedly delayed to 2028 on manufacturing snags, per SemiAnalysis reporting picked up by CNBC, and the company tightened its export-control compliance. TSMC is raising its US investment to 265 billion dollars. Databricks just got valued at 188 billion. Meanwhile regulators are circling too, the US is exploring a FINRA-style AI watchdog and the EU fined AliExpress 550 million euros under the DSA. Follow the money in every one of those stories and it lands on the same square: whoever controls fabrication and deployment capacity controls the industry. Model architecture is public research. Compute is the moat.
The valuations assume the pace holds
A skeptical note before anyone gets swept up. Moonshot is reportedly pursuing a Hong Kong IPO within about six months, and a new funding round could value it above 30 billion dollars, up from 20 billion in May. Its annualized revenue hit 300 million dollars in June, up from 200 million in April. That growth is real and fast. But 30 billion on 300 million in ARR is a 100x multiple, and it is a bet that this exact pace continues while the company literally cannot take new customers and its chip supply is a geopolitical football. Maybe it works out. I would just point out that the bull case requires the bottleneck to disappear, and nothing this week suggests it will.
The view from the small end of the market
I build iOS apps as a one-person studio, and my AI features run on-device, on the chip already in your pocket. So here is my honest read from the cheap seats: a 2.8 trillion parameter model changes almost nothing about my week. I cannot run it, I would not want to pay to serve it, and my users do not want their data leaving their phones to reach it anyway. What matters at my end of the market is what fits in a few gigabytes of RAM on an iPhone.
But the trickle-down is real, and it is the part I actually care about. Every frontier release becomes training data, distillation targets, and public technique for the next generation of small models. The 3 billion parameter models you can run on a phone today exist because the trillion-parameter ones existed first. So when two Chinese labs give away frontier weights in one week, the practical effect for someone like me is that on-device models get better and cheaper on a faster clock. The giants fight over compute. The rest of us quietly inherit the research.
If you want to see what on-device AI looks like in practice, private by default, nothing leaving your phone, the apps are at /apps.
Comments
Be kind and stay on topic. Comments are reviewed before they appear.