share_log

Zhipu AI claims its mysterious model "Niu Lai" runs on 100,000 domestically produced chips; SemiAnalysis comments that NVIDIA's moat is being tested once again.

wallstreetcn ·  Aug 27 09:30

Zhipu AI has officially claimed the mysterious model "Niu Lai" (Ox-Alpha) as its GLM-5.3-Flash. During testing, the model processed a daily traffic volume of 100 trillion tokens, entirely supported by over 100,000 domestic chips. Zhipu AI stated that its cost efficiency is now comparable to NVIDIA GPUs. Rumors suggest the chip supply may come from Huawei, Moore Threads, and Hygon. Furthermore, the model is priced at 1/40th of Claude Opus 4.8.

On August 26, Zhipu officially confirmed that the anonymous model Ox Alpha (known as "Niu Lai" in the Chinese community), which had sparked heated discussion in developer circles, is its newly released GLM-5.3-Flash. The model's codename was inspired by the recent domestic hit film "Niu Lai."

Prior to its official debut, this model was available for free testing in an anonymous capacity on OpenRouter and OpenCode. Within five days, cumulative traffic exceeded 50 trillion tokens, setting new traffic growth records for both platforms. Current usage is more than double that of DeepSeek. However, what truly captured market attention was a detail subsequently disclosed by Zhipu: all this traffic was handled entirely by domestically produced chips.

At the opening of Thursday's Hong Kong stock market session, $Z.AI (02513.HK)$ shares surged over 7% in response, closing at HK$1,106; $MINIMAX-W (00100.HK)$ followed by a 6% gain, after the company announced its first interim results since listing yesterday.

In the first half of the year, MINIMAX achieved revenue of USD 117 million, a year-on-year increase of 283.1%, surpassing the full-year 2025 revenue scale of USD 79 million. The revenue structure, previously dominated by AI-native products such as Hailuo AI, is shifting towards open platforms and enterprise services. In the first half, revenue from the open platform and other AI-based enterprise services reached USD 73.93 million, a year-on-year increase of 703.1%, with its share of total revenue rising to 63.4%, making it the company's largest revenue source. Revenue from AI-native products increased by 100.9% year-on-year to USD 42.64 million.

In its official technical documentation, Zhipu stated that a cluster of over 100,000 domestically produced chips was deployed for the model's inference services, noting that "hardware efficiency and cost per token have reached levels comparable to mainstream NVIDIA GPUs."

Semiconductor research firm SemiAnalysis subsequently commented on platform X: "All traffic was handled by domestically produced chips, with hardware efficiency and cost per token comparable to NVIDIA GPUs. Following yesterday's announcement regarding Jalapeño (OpenAI's self-developed inference chip), the CUDA moat is once again being tested."

100,000 Domestic Chips: From Testing to Production

Zhipu described the deployment details of this domestic chip cluster in its technical documentation.

The primary bottleneck for individual chips lies in memory capacity and bandwidth, making it particularly challenging to support context lengths of up to 1 million tokens. To address this, Zhipu has built a dedicated inference engine based on SGLang, employing techniques such as W8A8 quantization, INT8/FP8/BF16 mixed-cache quantization, and intra-node tensor parallelism. It also introduces a production-grade Encode–Prefill–Decode (EPD) decoupled architecture, separating multimodal encoding, prompt prefilling, and token-by-token decoding into independently schedulable work pools.

Zhipu states that end-to-end service performance has improved threefold compared to the initial baseline on the same hardware.

According to LatePost, the suppliers of these chips may include Huawei, Moore Threads, and Hygon. Zhipu’s official representatives declined to comment, and the specific chip models were not disclosed in technical documentation.

SemiAnalysis specifically noted in its commentary that while it was previously believed that only top-tier frontier laboratories possessed the computational scale to process 100 trillion tokens per day, "here, 100 trillion tokens of free traffic per day are entirely running on domestically produced chips."

Regarding the model itself: benchmarked against DeepSeek, with pricing at 1/40th that of Claude Opus 4.8.

GLM-5.3-Flash has a total parameter count of 320 billion, with 18 billion activated parameters. This represents approximately 40% of the parameter count of the previous-generation flagship GLM-5.3, nearly on par with DeepSeek V4 Flash.

In terms of performance, official evaluations by Zhipu show that GLM-5.3-Flash achieved a score of 57 on the Artificial Analysis Intelligence Index (AAI), surpassing the previous-generation flagship GLM-5.2, matching Anthropic's Claude Opus 4.8, and exceeding the score of 53 achieved by the official release of DeepSeek's flagship model V4 Pro.

Regarding pricing, GLM-5.3-Flash charges RMB 0.8 per million input tokens, RMB 2.8 per million output tokens, and RMB 0.23 for cache hits, which is one-tenth the price of GLM-5.3. A half-price discount will be applied during the first two weeks after launch.

Zhipu stated that the pricing for GLM-5.3-Flash is 1/10th that of GLM-5.3, 1/20th during the limited-time discount period, and 1/40th that of Claude Opus 4.8.

Compared to the adjusted pricing of DeepSeek V4 Flash (RMB 1.5 for input and RMB 4.5 for output during off-peak hours), GLM-5.3-Flash is more cost-effective in the vast majority of standard usage scenarios.

Architectural Innovation: Activated parameters and layer count nearly halved

GLM-5.3-Flash adopts a fundamentally different architectural design from GLM-5.3, which is the core reason it achieves higher performance at a lower cost.

Compared to GLM-4.5, GLM-5.3-Flash has a similar total parameter count (355B vs. 320B), but the number of activated parameters has decreased from 32B to 18B, and the number of layers has been reduced from 92 to 45, representing a near 50% reduction.

Zhipu AI states that GLM-5.3-Flash is the first open-source frontier model to employ a hybrid architecture combining sparse attention and linear attention. Compared to GLM-5.3, its attention computation volume and KV cache size are reduced by factors of 3.01 and 4.44, respectively.

Furthermore, this model is the first native multimodal model in the GLM-5 series, supporting image and video inputs. It also marks Zhipu AI's first launch of a new model with multimodal capabilities since strategically focusing on coding.

Breaking Through with Low Prices Amidst a Wave of Price Hikes

The release of GLM-5.3-Flash comes in the wake of a round of collective price increases in the Chinese AI model market.

Kimi K3's output price reached RMB 100 per million tokens, more than three times that of its predecessor, K2.6; Zhipu AI's output price rose to RMB 28 after its mid-year hike; DeepSeek V4 Pro increased from RMB 6 to RMB 13.5, reaching RMB 27 during peak hours; and DeepSeek V4 Flash's output price rose from RMB 2 to RMB 4.5, hitting RMB 9 during peak hours.

The direct consequence of these price hikes has been demand spillover. LatePost noted that following DeepSeek V4 Flash's price increase, its call volume on the OpenCode platform dropped by half, creating new growth opportunities for domestic model providers.

GLM-5.3-Flash's pricing strategy is specifically targeted at filling this gap.

Editor/rice

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.