share_log

Weekend Reading | Zhipu and MiniMax: Two Distinct Business Models for Large Language Models

Geekpark News ·  Sep 5 15:31

Source: GeekPark

Two major Chinese large language model companies released their first interim reports since listing within a five-day period.

$MINIMAX-W (00100.HK)$ Initially reported: H1 revenue reached USD 117 million, a year-on-year increase of 283%; August ARR exceeded USD 800 million. $Z.AI (02513.HK)$ Subsequently announced: H1 revenue amounted to RMB 954 million, a year-on-year increase of nearly 400%; August ARR reached USD 1.6 billion.

If focusing solely on ARR, the conclusion appears straightforward: Zhipu AI ranks higher, with MiniMax close behind; both companies cite surging token usage, growth in enterprise clients, and the initial monetization of model capabilities.

A stratification in the commercialization of China's large language models.

Reconciling these figures with financial statements reveals a more complex picture.

Zhipu AI recognized approximately USD 142 million in revenue during the first half, while MiniMax recorded USD 117 million; their actual revenue scales are relatively close. Both companies also reported similar adjusted losses, each nearing USD 300 million. While ARR reflects the revenue run rate at a specific point in time, it cannot directly substitute for revenue, gross profit, or cash flow.

Neither company has yet achieved profitability. In both interim reports, the 'narrowing of losses for the period' was influenced by accounting effects stemming from the conversion of preferred shares post-listing. Stripping away these factors, R&D and operational expenditures remain at elevated levels.

The key difference lies in where each company is directing its model sales.

API revenue accounts for 86.5% of Zhipu AI's total income, as it begins to deploy its models into coding, cybersecurity, and long-horizon task environments. Meanwhile, MiniMax has seen rapid growth in enterprise service revenue while maintaining its core focus on globalization, multimodality, and AI-native products. It prioritizes reducing the cost per unit of intelligence to enable broader consumption by users, agents, and creative scenarios.

These two interim reports illustrate a stratification in the commercialization of China's large language models.

1. Zhipu AI: The tide of on-premise deployment is receding; APIs have become the primary growth engine.

Let us first examine Zhipu AI's interim report. Revenue for the first half of the year reached RMB 954 million, a year-on-year increase of 399.7%, surpassing the full-year 2025 revenue of RMB 724 million. Gross profit stood at RMB 252 million, with a gross margin of 26.4%. R&D expenditure amounted to RMB 213.1 million, representing a year-on-year increase of 33.6%.

Image source: Zhipu AI Interim Report

Behind this growth lies a significant shift in business structure. Revenue from the open platform and API businesses reached RMB 825 million, a year-on-year surge of 2,735.7%, accounting for 86.5% of total revenue—compared to just 15.2% in the same period last year. On-premise deployment, which previously formed the core business base, has moved to a secondary position: revenue related to enterprise general-purpose models fell to RMB 67.04 million, a year-on-year decline of 54.6%.

Zhipu AI is transitioning from a model of "deploying models in customer data centers and recognizing revenue upon project completion" to one of "hosting models on the cloud and charging continuously based on usage volume." The former resembles traditional project-based business, emphasizing delivery, security, and client relationships; the latter relies on model capabilities, developer ecosystems, and inference costs, offering greater potential for sustainable and scalable revenue.

As of the release of the interim report, token consumption on the MaaS platform had increased more than 40-fold since the beginning of the year. Daily active paying users grew by 603%, and the average selling price (ASP) of APIs rose by approximately 101%. The gross margin for the open platform and API businesses also improved significantly, rising from -0.4% in the same period last year to 24.6%.

During the conference call, management stated that annualizing August's monthly revenue implies an Annual Recurring Revenue (ARR) of USD 1.6 billion for the MaaS platform; annualizing the most recent week's revenue suggests an ARR exceeding USD 2 billion. Although ARR is not included in financial statements and short-term annualization can exaggerate monthly growth rates, Zhipu AI's emphasis on this metric reflects its desire for the market to evaluate the company based on the "current operational velocity of its cloud-based model services," rather than solely on recognized semi-annual revenue.

The rapid growth in API revenue has not altered the reality that Zhipu AI remains in a phase of heavy investment. As the revenue structure shifted from high-margin on-premise deployments to APIs, the overall gross margin declined from 50.0% in the same period last year to 26.4%. The company reported a loss of RMB 2.072 billion during the period, with an adjusted net loss of RMB 1.964 billion. Operating losses widened from RMB 1.899 billion to RMB 2.147 billion. The divergence in loss figures is primarily due to the conversion of preferred shares: fair value changes in related financial liabilities resulted in a loss of RMB 429 million in the same period last year, which decreased to RMB 22.09 million this year. However, operating losses, which better reflect core business performance, still expanded from RMB 1.899 billion to RMB 2.147 billion. R&D investment during the period reached RMB 213.1 million, more than twice the revenue, indicating that substantial capital is still required for model training and inference infrastructure.

However, Zhipu AI did not rely solely on traditional financial metrics to justify this investment. Its 60-page interim results announcement devoted considerable space to discussing model architecture, training algorithms, reinforcement learning, inference systems, and domestic chip clusters, concluding with a glossary of technical terms.

In the announcement, Zhipu AI proposed a formula: "Commercial Value of AGI = Upper Bound of Intelligence × Scale of Token Consumption." The logic is that the upper bound of intelligence determines the tasks a model can handle, the task boundaries determine the value density per token, and the scale of token consumption amplifies this value. Tokens used for casual chat represent different commercial value compared to those used for completing software engineering tasks or identifying vulnerabilities in production systems. The company also defined a custom "computing power multiplier" to measure how much open platform and API revenue is generated per yuan invested in training and inference computing power, claiming it increased approximately 14-fold year-on-year in the first half—although absolute values and calculation methods were not disclosed, preventing independent verification by external parties.

By incorporating these custom metrics into its interim report, Zhipu AI is seeking to redefine the valuation narrative for large language model (LLM) companies. Beyond revenue, profit, and cash flow, it aims to direct investor attention to model capabilities, token consumption, cost per task unit, and the efficiency of converting computing power into revenue.

Recent developments in the model can be understood through this methodological framework.

The GLM-5.2, released in June, focused on long-context coding and agent capabilities. The subsequent GLM-5.3 retained the same foundational architecture and parameter scale but primarily expanded the environment for long-horizon tasks and increased investment in reinforcement learning. Its end-to-end task completion rate in self-built real-world coding benchmarks improved by over 50% compared to version 5.2, reflecting an attempt to raise the ceiling for the model’s ability to handle complex tasks independently.

Launched in late August, GLM-5.3-Flash is designed to reduce the cost of large-scale inference. Featuring a sparse architecture with 320 billion total parameters and 18 billion activated parameters, it operates on a cluster of approximately 100,000 domestic chips. Priced at roughly one-tenth that of GLM-5.2, it consumed over 62 trillion tokens during six days of anonymous testing. Zhipu AI also disclosed that the inference cost per token has decreased by approximately 80% since the beginning of the year.

GLM-5.3 expands the upper limit of task complexity, while GLM-5.3-Flash scales up token consumption; these two product lines correspond to the two ends of the valuation equation.

The next challenge is to demonstrate whether these capabilities can extend further into deeper workflow integration.

Currently, revenue from the agent business stands at RMB 55.56 million, representing a year-on-year increase of 304.4%. While growth is rapid, the base remains limited. Cybersecurity has entered the co-work validation phase, whereas applications in legal, financial, and educational sectors are still in early stages. "Selling outcomes" is currently more of a pathway under construction: evolving from selling localized models to offering API calls and coding subscriptions, then integrating agents into enterprise processes, and ultimately charging based on end-to-end task completion. The rapid growth in API usage validates one segment of this strategy. Whether co-work solutions will become the next major revenue stream remains to be seen in future financial reports.

02. MiniMax: More Globalized Revenue, Prioritizing Affordable "Unit Intelligence"

MiniMax's interim report presents a different growth profile.

In the first half of 2026, MiniMax generated revenue of USD 117 million, a year-on-year increase of 283.1%; gross profit reached USD 20.81 million, up 464.8% year-on-year, with the gross margin rising from 12.1% in the same period last year to 17.9%. Of this, revenue from the open platform and other AI enterprise services amounted to USD 73.93 million, surging 703.1% year-on-year and accounting for 63.4% of total revenue. Revenue from AI-native products was USD 42.64 million, representing a 100.9% year-on-year increase.

Image source: MiniMax interim report

Compared with Zhipu, one distinct difference for MiniMax stands out: overseas markets have already become the primary revenue source. In the first half of the year, MiniMax’s overseas revenue amounted to USD 70.83 million, accounting for 60.8% of total revenue; revenue from mainland China was USD 45.75 million, representing 39.2%. During the conference call, the company further disclosed that its enterprise and developer customer base exceeded 2 million, a nearly tenfold increase from the end of last year. Its August Annualized Recurring Revenue (ARR) surpassed USD 800 million, with the proportion of To-B business in ARR rising from approximately 30% a year ago to around 80%.

These are also figures reflecting rapid expansion. MiniMax did not fully explain in its financial report how the USD 800 million ARR was annualized, so it cannot be directly subtracted from Zhipu’s USD 1.6 billion figure, which was calculated by multiplying August revenue by 12. The two companies also differ in their definitions of customers, revenue recognition policies, business composition, and ARR calculation methodologies.

Nevertheless, the operational logic presented by MiniMax during the conference call was clear: model prices can continue to decline, while gross profit per unit of Token can still improve. This is driven by joint optimizations in model architecture, computing power scheduling, supply chain, and inference efficiency. The company summarizes this approach as “Minimize the Cost, Maximize the Intelligence”—aiming to reduce the cost of achieving equivalent intelligence as much as possible, given that computing power and energy remain scarce resources.

This logic is also reflected in MiniMax’s recent product rollout pace.

Launched in June, MiniMax-M3 focuses on coding, agentic capabilities, and million-level context windows. During the earnings call, the company stated that M3 delivers stronger capabilities than its predecessor at a similar price point, which has both attracted new customers and encouraged existing clients to increase their usage.

The subsequently released H3 focuses on native multimodal generation, such as video and audio, and made its weights open-source in August. For MiniMax, text models and multimodal models are not two unrelated product lines: the former addresses enterprise development, coding, and agent-related needs, while the latter targets creation, content production, and a broader range of AI-native applications.

MiniMax is also adapting to domestic chips and has announced that M3.1 will further reduce inference costs. Management mentioned that the goal is to lower inference costs to approximately one-third of the initial level of M3. While this target could easily be interpreted as a price war, MiniMax is more concerned with whether cost reductions can drive greater Token consumption, attract more developers, and expand commercial coverage.

Its financial report also reminds the market that there is still a distance between growth and profitability. In the first half of the year, MiniMax’s R&D expenses were USD 297 million, a year-on-year increase of 138.8%; the net loss for the period was USD 358 million, narrowing by approximately 11% year-on-year. However, the adjusted net loss widened by 111.2% year-on-year to USD 293 million. The apparent improvement in the net loss was also influenced by a reduction in changes to financial liabilities related to preferred shares after listing: the fair value loss from this item decreased from USD 254 million in the same period last year to USD 31 million.

MiniMax’s gross margin is indeed improving, and its revenue is shifting towards enterprise clients and overseas markets; however, the company is still bearing high costs for model training, multimodal products, and global channel expansion. Its bet is not on immediately raising the price per API call, but rather on making each unit of intelligence more efficient, thereby making it affordable and accessible to more users, which in turn drives higher usage.

3. Following their earnings releases, the two companies are addressing different questions.

Viewing the two financial reports side by side, Zhipu and MiniMax both demonstrate the same point: model capabilities can be converted into revenue, and token consumption is becoming a trackable operational metric. The differences between the two companies lie primarily in the sequencing of their commercialization strategies.

Zhipu is driving its billing units up the value chain. Previously, customers purchased models deployed locally; now, they pay based on usage volume, with subscriptions emerging as the new transaction model. The company aims to advance further into co-work scenarios, integrating models into software engineering, cybersecurity, and knowledge-intensive work, ultimately charging per task, process, or deliverable outcome. This approach presupposes that as models cross each capability threshold, they can handle higher-value tasks—thereby gradually decoupling pricing from token counts and aligning it more closely with the actual work completed by the model.

MiniMax, meanwhile, focuses more on expanding the reach of its intelligence. It chooses to lower the unit cost of intelligence through architectural and efficiency optimizations, leveraging multimodal capabilities across text, video, and audio to embed AI into a broader range of developer, enterprise, and content scenarios. Price reductions drive higher usage volumes, which in turn feed back into cost optimization, while overseas revenue accounting for 60% of the total provides a sufficiently large market radius for this scale-driven logic.

These two paths will eventually converge at the same point: For Zhipu to persuade enterprises to pay for outcomes, it must reduce the cost of task completion to a level suitable for scalable adoption; for MiniMax to generate higher value from low-cost intelligence, its models must reliably execute complex tasks. Capabilities determine how deeply AI can penetrate work processes, while costs determine how broadly these capabilities can reach users. Both elements are indispensable.

Consequently, the questions to be addressed in the next earnings report are more specific: When will Zhipu’s co-work model become a significant source of revenue? Can MiniMax’s declining unit intelligence costs continue to attract more overseas customers and sustain healthier gross margins?

One company is pushing AI into deeper work processes, while the other is expanding AI into broader markets. Only when models can both complete tasks effectively and become affordable enough for mass adoption will large language models truly overcome the most difficult hurdle in commercialization.

Editor/rice

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.