The new intelligent model "Niu Lai" runs on a domestically produced graphics card costing 100,000 yuan, prompting SemiAnalaysis to comment that "Nvidia's competitive advantage is being tested again."

The new intelligent model "Niu Lai" runs on a domestically produced graphics card costing 100,000 yuan, prompting SemiAnalaysis to comment that "Nvidia's competitive advantage is being tested again."

On August 26, Zhipu officially confirmed that the anonymous model Ox Alpha (known as "Niu Lai" in the Chinese community), which had previously sparked heated discussions in the developer community, is its latest GLM-5.3-Flash. The model's codename was inspired by the recently released Chinese film "Niu Lai".

Before its official launch, this model was anonymously and freely tested on OpenRouter and OpenCode, accumulating over 50 trillion tokens of traffic within five days, breaking the traffic growth records of both platforms. Its current usage is more than double that of DeepSeek. However, what truly captured market attention was a detail subsequently disclosed by Zhipu: all this traffic was carried by domestically produced chips.

In its official technical documentation, Zhipu stated that it utilizes a cluster of over 100,000 domestically produced chips for the model's inference service , and that "the hardware efficiency and cost per token have reached a level comparable to mainstream NVIDIA GPUs."

Semiconductor research firm SemiAnalysis subsequently commented on the X platform: " All traffic is carried by domestically produced chips, and the hardware efficiency and cost per token are comparable to NVIDIA GPUs. Following the announcement of Jalapeño (OpenAI's self-developed inference chip) yesterday, CUDA's competitive advantage is once again being tested."

100,000 domestically produced cards: From testing to production

Zhipu described the deployment details of this domestically produced chip cluster in its technical documentation.

The main bottlenecks of a single chip are memory capacity and bandwidth, making it particularly difficult to support context lengths of up to 1 million tokens. To address this, Zhipu built a dedicated inference engine on top of SGLang, employing technologies such as W8A8 quantization, INT8/FP8/BF16 hybrid cached quantization, and intra-node tensor parallelism. It also introduced a production-grade Encode–Prefill–Decode (EPD) split architecture, separating multimodal encoding, prompt prefilling, and token-by-token decoding into independently schedulable work pools.

According to Zhipu, the end-to-end service performance was improved by 3 times compared to the initial baseline on the same hardware.

According to LatePost, the suppliers of these chips may be Huawei, Moore Threads, and Hygon. Zhipu has not commented on this, and its technical documents do not specify the exact chip model.

SemiAnalysis specifically pointed out in its commentary that some previously held the view that only top-tier cutting-edge laboratories possessed the computing power to process 100 trillion tokens per day, "while here, the 100 trillion tokens of free traffic per day all run on domestically produced chips."

The model itself: comparable to DeepSeek, priced at 1/40th of Claude Opus 4.8.

The GLM-5.3-Flash has a total of 320B parameters, with 18B active parameters. The number of parameters is about 40% of the previous generation flagship GLM-5.3, and almost the same as DeepSeek V4 Flash.

In terms of performance, according to the official evaluation of Zhipu, the GLM-5.3-Flash scored 57 points in the Artificial Analysis Intelligence Index (AA), which is higher than the previous generation flagship GLM-5.2, on par with Anthropic's Claude Opus 4.8, and higher than the 53 points of the official version of DeepSeek's flagship model V4 Pro.

In terms of pricing, GLM-5.3-Flash costs 0.8 yuan per million token inputs and 2.8 yuan per million token outputs, with a cache hit price of 0.23 yuan, which is one-tenth of that of GLM-5.3. It will be half-price for the first two weeks after launch.

Zhipu stated that the GLM-5.3-Flash is priced at 1/10 of the GLM-5.3, and during the limited-time discount, it is 1/20 of the GLM-5.3 and 1/40 of the Claude Opus 4.8.

Compared to the adjusted DeepSeek V4 Flash (costing 1.5 yuan for input and 4.5 yuan for output during off-peak hours), GLM-5.3-Flash is cheaper in most common usage scenarios.

Architectural innovation: activation parameters and number of layers are nearly halved.

GLM-5.3-Flash adopts a completely different architecture design from GLM-5.3, which is the core reason why it can achieve higher performance at a lower cost.

Compared to GLM-4.5, GLM-5.3-Flash has a similar total number of parameters (355B vs. 320B), but the number of active parameters is reduced from 32B to 18B, and the number of layers is reduced from 92 to 45, almost halved.

According to Zhipu, GLM-5.3-Flash is the first open-source cutting-edge model to adopt a hybrid architecture of sparse attention and linear attention. Compared to GLM-5.3, its attention computation and key-value cache size are reduced by 3.01 times and 4.44 times, respectively.

Furthermore, this model is the first native multimodal model in the GLM-5 series, supporting image and video inputs. It is also the first new model with multimodal capabilities launched by SmartSpectrum since it focused on coding.

Breaking through price barriers amidst a wave of price increases

The release of GLM-5.3-Flash occurred after a round of collective price increases in the Chinese AI model market.

Kimi K3's output price reached 100 yuan per million tokens, more than three times that of its predecessor, K2.6; Zhipu's output price also reached 28 yuan after raising prices in the first half of the year; DeepSeek V4 Pro rose from 6 yuan to 13.5 yuan, reaching 27 yuan during peak periods; DeepSeek V4 Flash's output price rose from 2 yuan to 4.5 yuan, reaching 9 yuan during peak periods.

The direct consequence of the price increase is an outflow of demand. LatePost points out that after the price increase of DeepSeek V4 Flash, the number of calls on the OpenCode platform dropped by half, and this demand has become a new growth area for domestic model manufacturers.

The pricing strategy of GLM-5.3-Flash is precisely aimed at this gap.

Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.