Ranked ninth globally and first among open-source models, Kimi K3’s breakthrough was immediately met with a “compute freeze”—ironically confirming the logic of hardware demand.

Ranked ninth globally and first among open-source models, Kimi K3’s breakthrough was immediately met with a “compute freeze”—ironically confirming the logic of hardware demand.

```

Open-source large models have never been so close to world-class levels—yet never before has their launch so dramatically demonstrated that the demand for computing power is far from its peak.

On July 17, Moonshot AI released its new flagship model, Kimi K3. On the authoritative Text Arena comprehensive ranking, this model ranked 9th with a score of 1486, becoming the first open-source model to enter the global top ten.

Even more notably, on the Frontend Code Arena leaderboard, Kimi K3 surpassed Anthropic’s Claude Fable 5 with 1679 points to take first place, ranking first in 6 out of 7 frontend code subcategories.

However, within 48 hours after its release, due to user request volumes greatly exceeding estimates and approaching cluster capacity limits, Moonshot AI had to announce on July 19 that subscriptions for new consumer users would be suspended—the launch of an open-source model actually ended due to “insufficient computing power.”

This seemingly paradoxical phenomenon precisely reveals the core debate in today’s AI industry: Do more efficient models reduce or amplify hardware demand?

From 38th to 9th: Open Source Breaks into the Top Ten for the First Time

The leap in Kimi K3’s ranking is remarkable. The previous generation Kimi K2.6 was only 38th on Text Arena, while K3 jumped directly to 9th place.

According to China Merchants Securities Strategy Research’s latest weekly report, the previous highest-ranking open-source model was Qwen 3.7-max-preview at 18th, making the emergence of K3 the first time an open-source model has entered the top tier alongside the world’s leading closed-source models in a comprehensive text evaluation.

Looking at each category, K3 ranked in the top ten for creative writing, code generation, and instruction following, and was first in specialized fields such as physical and social sciences, legal and government, and healthcare; it ranked third in business management and financial operations.

Technically, K3 has 2.8 trillion parameters, adopts Kimi Delta Attention and Attention Residuals architecture, and uses Stable LatentMoE to achieve high sparsity—activating only 16 out of 896 experts per token. According to official estimates, the overall architecture is about 2.5 times more scalable than Kimi K2.

At the Top, Yet “Circuit Breaker”: Reverse Evidence of Computing Bottleneck

The most dramatic moment after K3’s release was not its ranking—but the subscription “circuit breaker.” In an announcement, Kimi detailed that in the last 48 hours, user requests almost hit cluster limits, so consumer sign-ups are suspended effective immediately while the company focuses on scaling up computing.

This directly contradicts Goldman Sachs’ earlier prediction. Rich Privorotsky, a partner at Goldman Sachs, had pointed out that a Chinese lab—without matching the largest-scale Western pretraining compute—had narrowed the gap with the top US models through architectural innovation, using this to warn that “scaling is no longer the only path to victory.”

But Kimi’s subscription “circuit breaker” provides reverse evidence—when model performance is truly recognized by the market, the shortfall in computing power appears in an even more dramatic way.

A report from Meritz Securities Korea on July 19 also directly addressed this debate. The report distinguishes K3 from DeepSeek: DeepSeek’s core narrative is extremely low $6-million training cost, while K3 does not disclose its training cost and officially recommends deployment on a hypernode environment with at least 64 high-performance chips.

In terms of usage cost, each K3 task costs $0.95—similar in scale to GPT-5.6 Sol’s $1.04 and Claude Fable 5’s $2.75.

Closed-Source Moats Narrow, but “Jevons Paradox” is at Work

China Merchants Securities’ research further points out that the most direct impact of open-source models on closed source comes from the performance catch-up.

According to the Stanford AI Index 2026, as of March, top closed-source models lead top open-source ones by about 3.3%—a gap no longer big enough to justify closed-source maintaining premium pricing in all scenarios. Research by Epoch AI shows that the strongest open-weight models on average lag behind the strongest closed-source ones by roughly 4 months.

Looking at the trend in token consumption, the Silicon Data LLM Token Expenditure Index has been declining since June, indicating AI services are following a cloud-computing-like trend: falling costs drive expanded usage scale.

According to OpenRouter data, global weekly invocation had already reached 52.6 trillion tokens as of July 6; open-source models like DeepSeek and Qwen have repeatedly led global usage thanks to their cost-effectiveness.

Citi analysts view K3 as a potential case of the “Jevons Paradox” in the AI industry chain: When technical efficiency reduces unit cost, overall usage may surge, ultimately driving resource consumption higher rather than lower.

BofA Securities also points out that if Chinese open-source models continue to catch up, US frontier AI labs are more likely to increase—rather than decrease—computing investment to maintain differentiation.

Rise of Hypernodes: Entering System-Level Competition

As demand for computing power keeps rising and the capability of a single chip can no longer alone determine system strength, hypernodes are becoming the core direction for AI infrastructure evolution.

China Merchants Securities described in its weekly report that hypernodes are new high-density computing systems that use high-speed, low-latency scale-up interconnects to link hundreds or even thousands of AI accelerator chips into a unified logical compute unit.

At the 2026 WAIC, Huawei Ascend 950 hypernode hardware was unveiled for the first time, supporting up to 1024-card deployments and providing 1 EFLOPS FP8 compute and 256TB unified memory addressing space.

BCC Research predicts the global data center networking technology market will grow from $45.8 billion in 2025 to $103 billion in 2030, a compound annual growth rate of about 17.6%.

Research and Markets expects the AI high-speed interconnect market to grow from $10.31 billion in 2025 to $30.17 billion in 2030, a compound growth rate of 23.9%—much higher than traditional data center networking growth.

For investors, the release of Kimi K3 sends a clear signal: breakthroughs in open-source models will not weaken the underlying hardware demand. The additional workload unleashed by efficiency improvements is shifting industry competition from the models themselves to the computing infrastructure that supports them.

Risk Warning and DisclaimerThe market has risks; investments require caution. This article does not constitute personal investment advice and does not take into account individual users’ specific investment goals, financial situations, or needs. Users should consider whether any opinions, viewpoints, or conclusions in this article are suitable for their specific situation. Investment decisions made accordingly are at one’s own risk. ```