With over double the call volume of DeepSeek and a score of 57 comparable to Opus 4.8, Zhipu's stock price surged over 9% intraday.

With over double the call volume of DeepSeek and a score of 57 comparable to Opus 4.8, Zhipu's stock price surged over 9% intraday.

The unveiling of Zhipu's mysterious Ox Alpha model has driven the company's stock price to a significant rise. The disclosure that it can handle massive inference traffic entirely with domestically produced chips has also attracted widespread attention in the AI hardware field.

On Thursday, Zhipu's stock price surged over 9% intraday, reaching HK$1124. In related news, Zhipu officially confirmed on August 26th that its newly released GLM-5.3-Flash (320B-A18B) is the anonymous model Ox Alpha—known as "Niu Lai" in the Chinese community—which had previously generated considerable buzz in the developer community. This model scored 57 points in the Artificial Analysis Intelligence comprehensive intelligence index, matching Anthropic's Claude Opus 4.8 and surpassing the 53 points scored by the official version of DeepSeek's flagship model V4 Pro.

Prior to its official release, GLM-5.3-Flash was offered for free and anonymous testing on the OpenRouter and OpenCode platforms. Within five days, it accumulated over 50 trillion tokens in traffic, breaking traffic growth records on both platforms. Its current usage is more than double that of DeepSeek. Zhipu also disclosed that all of this traffic was handled by domestically produced chips, utilizing over 100,000 domestically produced chip clusters for inference services, and stated that "hardware efficiency and per-token cost have reached levels comparable to mainstream NVIDIA GPUs."

Domestically produced chips are fully integrated, and the hardware narrative has attracted market attention.

In its latest disclosure, the performance of domestically produced chips was one of the key focuses of market attention. According to Zhipu's official technical documentation, the company has deployed over 100,000 domestically produced chips to form a cluster, providing inference computing power support for GLM-5.3-Flash, and stated that its hardware efficiency and cost per token are comparable to NVIDIA GPUs.

Semiconductor research firm SemiAnalysis subsequently commented on the X platform, stating, "All traffic is carried by domestically produced chips, and the hardware efficiency and cost per token are comparable to NVIDIA GPUs. Following yesterday's announcement of Jalapeño (OpenAI's self-developed inference chip), CUDA's competitive advantage is once again being tested." The firm also specifically pointed out that previously, the industry generally believed that only top-tier cutting-edge laboratories possessed the computing power to process 100 trillion tokens per day, "while here, the free traffic of 100 trillion tokens per day is all running on domestically produced chips."

According to LatePost, the suppliers of these chips may be Huawei, Moore Threads, and Hygon. Zhipu has not commented on this, and its technical documents do not specify the exact chip model.

To address the bottlenecks of single-chip memory capacity and bandwidth, Zhipu built a dedicated inference engine on top of SGLang, employing technologies such as W8A8 quantization, INT8/FP8/BF16 hybrid cache quantization, and intra-node tensor parallelism, and introducing a production-grade Encode–Prefill–Decode (EPD) separated architecture. Zhipu claims that compared to the initial baseline on the same hardware, end-to-end service performance has been improved by 3 times.

Architecture Restructuring: Activation Parameters and Number of Layers Nearly Halved

The GLM-5.3-Flash has significant differences in architecture compared to its predecessors, which is the core reason why it can achieve higher performance at a lower cost.

Compared to GLM-4.5, GLM-5.3 Flash has a similar total number of parameters (355B vs. 320B), but the number of active parameters has decreased from 32B to 18B, and the number of layers has been reduced from 92 to 45, almost halved. The number of parameters is about 40% of the previous generation flagship GLM-5.3, and almost the same as DeepSeek V4 Flash.

According to Zhipu, GLM-5.3-Flash is the first open-source cutting-edge model to adopt a hybrid architecture of sparse attention and linear attention. Compared to GLM-5.3, its attention computation and key-value cache size are reduced by 3.01 times and 4.44 times, respectively. Furthermore, this model is the first native multimodal model in the GLM-5 series, supporting both image and video inputs. It is also the first new model with multimodal capabilities launched by Zhipu since its strategic focus on coding.

Pricing strategy: Targeting the demand gap following the price increase trend

The release of GLM-5.3-Flash coincided with a round of collective price increases in the Chinese AI model market. The output price of Kimi K3 reached 100 yuan per million tokens, more than three times that of its predecessor, K2.6; DeepSeek V4 Pro rose from 6 yuan to 13.5 yuan, reaching 27 yuan during peak periods; and the output price of DeepSeek V4 Flash rose from 2 yuan to 4.5 yuan, reaching 9 yuan during peak periods.

According to LatePost, after DeepSeek V4 Flash raised its price, its usage on the OpenCode platform dropped by half, and this spillover demand has become a new growth area for domestic model manufacturers.

GLM-5.3-Flash's pricing strategy targets this gap. The model costs 0.8 yuan per million token inputs and 2.8 yuan per million tokens output, with a cache hit price of 0.23 yuan, one-tenth that of GLM-5.3. It will be half-price for the first two weeks after launch. Zhipu stated that during the limited-time discount period, the price will be 1/20th of GLM-5.3 and 1/40th of Claude Opus 4.8. Compared to the adjusted price of DeepSeek V4 Flash, GLM-5.3-Flash is cheaper in most common usage scenarios.

Zhipu plans to release its first detailed financial report next Monday, covering the first six months of its performance since its listing in January this year. The release of GLM-5.3-Flash and its market performance will provide investors with an important benchmark for assessing the company's commercialization progress.

Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.