Morgan Stanley recently held a closed-door meeting to conduct in-depth discussions on the latest developments among domestic AI large-model companies and major internet firms.
Morgan Stanley recently held a closed-door meeting to conduct in-depth discussions on the latest developments among domestic AI large-model companies and major internet firms. The meeting minutes indicate that MiniMax and Zhipu, two leading large-model companies, are both maintaining rapid revenue growth. MiniMax’s Annual Recurring Revenue (ARR) reached $800 million by the end of August, and analysts have raised their year-end forecast to $1.3 billion. Zhipu’s weeklyized ARR exceeded $2 billion in August, with its year-end guidance raised to $2.4 billion. Tencent’s Hunyuan 4 Preview was released at the end of August, with internal blind tests showing performance slightly superior to GLM-5.3 and Kimi K3. The WorkBody open platform is also evolving from an office assistant to an agent operation layer.
Analysts believe that buy-side investors are currently most concerned with model capabilities, followed by ARR growth and gross margins. Model competition will enter a phase of stratified elimination. Companies with State-of-the-Art (SOTA) capabilities can launch cost-effective versions through distillation, whereas companies relying solely on cost-effectiveness will struggle to break through technical barriers and face pressure on gross margins.
MiniMax: Year-end ARR expected to exceed expectations; three new models pending release
MiniMax’s ARR reached $800 million by the end of August. While the company maintains its guidance of over $1 billion for year-end, analysts have raised their expectation to $1.3 billion. The company plans to release three new models—M3.1, H3.1, and M3 Pro—in the second half of the year. Current computing power reserves are sufficient to support all training activities within the year and to prepare in advance for next year’s 10-trillion-parameter model.
Zhipu: Rapid ARR growth; domestic chip clusters support trillion-parameter training
Zhipu is experiencing rapid ARR growth, with its weeklyized ARR exceeding $2 billion in August, and its year-end guidance raised to $2.4 billion; the actual scale depends on the delivery progress of domestic chips. The GLM-5.3 Flash model released by the company in August adopts a new architecture and performs inference on a fully domestic chip cluster. This architecture will be used for GLM-6, scheduled for release in October. Zhipu possesses a domestic chip cluster with 100,000 cards. Its top 10 customers contribute over 40% of its revenue, and nine of the top 10 internet companies in China are its clients.
Tencent Hunyuan: Accelerated iteration; WorkBody builds an open Agent ecosystem
Tencent’s Hunyuan 4 Preview was released at the end of August, featuring significant expansions in parameter count and context length. Internal blind tests showed its performance to be slightly superior to GLM-5.3 and Kimi K3. WorkBody is positioned as an open Agent ecosystem connecting hardware, applications, and developers, and is accelerating the development of payment and distribution capabilities. If Hunyuan becomes a SOTA model, it can develop into a Model-as-a-Service offering; even if it does not achieve SOTA status, the WorkBody ecosystem retains commercial value. Future focus will be on new product and strategy announcements at the Global Digital Ecosystem Conference in October.
Appendix: Q&A Session
Q: What was MiniMax's Annual Recurring Revenue (ARR) as of the end of August? What is the statistical methodology?
A: As of the end of August, MiniMax's ARR stood at USD 800 million, calculated on a weekly basis. This methodology is consistent with the figure of over USD 400 million disclosed in May. The data exceeded market expectations.
Q: What is MiniMax's ARR guidance for year-end? How do market expectations compare with the company's explanation?
A: The company maintains its year-end ARR guidance at above USD 1 billion, although this guidance is considered conservative. Analysts have raised their expectations to USD 1.3 billion. The company has not revised its guidance upward as it awaits the actual performance following the release of new models in the second half of the year; the potential for an upward revision remains to be observed.
Q: Which new models does MiniMax plan to release in the second half of the year? What are the launch timelines and core features?
A: MiniMax plans to release three new models: M3.1, H3.1, and M3 Pro. The company currently has sufficient computing power reserves to support the training of all models within the year and is pre-positioning computing resources for a 10-trillion-parameter model next year.
Q: What were MiniMax's R&D expenses in the first half of the year? What are the expectations for R&D spending in the second half and next year?
A: R&D expenses in the first half amounted to approximately USD 300 million. Annualized R&D expenses for the second half are expected to exceed USD 1 billion. R&D spending will increase further next year, but the growth rate will be significantly lower than that of revenue and gross profit.
Q: What were the primary reasons for the decline in MiniMax's gross margin in the first half, and what is the outlook for the second half?
A: The main reasons for the gross margin decline include: price reductions and user subsidies resulting from poor initial market reception of the M3 model; the ramp-up period for inference optimization of new models; lower gross margins and significant discounts associated with token plans; and an increased share of revenue from text models, whereas multimodal models have gross margins exceeding 50%. The company noted that most factors in the first half were one-off events and expects quarter-on-quarter improvement in gross margin in the second half.
Q: How does MiniMax view the relationship between high-intelligence models and cost-effective models?
A: MiniMax believes that intelligence level and cost-effectiveness are two sides of the same coin. High inference efficiency and low costs enable larger-scale post-training, thereby enhancing model intelligence. Therefore, inference efficiency is a core capability that cannot be overlooked in the pursuit of high intelligence.
Q: What is the timing window for MiniMax's next financing round, and what are the company's plans?
A: The cooling-off period for the previous financing round ends on September 12, allowing the company to initiate fundraising as early as September 13. However, management has stated that it does not plan to raise funds in the short term, intending to proceed after releasing an industry state-of-the-art (SOTA) model. The market generally expects the financing to occur after October.
Q: What are Zhipu AI's ARR figures at various points in time, and what is the statistical methodology? What is the year-end guidance and market feedback?
A: Zhipu AI's ARR data: USD 250 million in March, USD 530–540 million in June, USD 1 billion in July, and USD 1.6 billion in August. On a weekly basis, August ARR exceeded USD 2 billion. The year-end guidance has been raised to USD 2.4 billion. The company considers this guidance conservative, noting that actual scale depends on computing power delivery progress. It is expected that guidance will be updated in the fourth quarter based on the delivery status of domestic chips.
Q: What are the key updates to Zhipu AI's model pipeline in the second half of the year? What is the positioning and characteristics of GLM-6?
A: GLM-5.3 Flash, released in August, adopts a new architecture and performs inference on clusters of fully domestic chips. This architecture will be used for GLM-6, scheduled for release in October. GLM-5.3 may be updated to version 5.5 through post-training; however, if GLM-6 training is completed, the company may skip version 5.5 and release GLM-6 directly.
Q: What were the reasons for the decline in Zhipu AI's gross margin in the first half of the year, and what is the long-term gross margin target?
A: The primary reasons for the decline in gross margin include significant fluctuations in computing costs, frequent model iterations leading to user switching and inference ramp-up periods, and lower efficiency during the initial optimization phase of domestic chip clusters. The company plans to improve inference efficiency through collaborative optimization across the model, network, and chip layers, aiming to raise the gross margin of its open platform to over 50% within the next 12–18 months.
Q: What is the status of Zhipu AI's computing power reserves and customer structure?
A: Zhipu AI possesses a cluster of 100,000 domestically produced chips and has reserved sufficient advanced computing power to train models with trillions of parameters. The top 10 customers contribute over 40% of revenue and Annual Recurring Revenue (ARR). Nine of the top 10 internet companies in China are Zhipu AI customers, with four of them adopting the GLM model as their primary national strategy.
Q: How does Zhipu AI view the market space for AI and its future expansion directions?
A: The company believes that the global coding market is valued at approximately $500 billion, with current penetration rates still low. The next step involves expanding from coding to coworking, a market space ten times larger than coding. In the long term, the Total Addressable Market (TAM) for AI is projected to reach $30 trillion. GLM-5.3 Flash has already demonstrated cybersecurity capabilities, supporting expansion towards Copilot functionalities.
Q: What is the financing window and market expectation for Zhipu AI?
A: The cooling-off period for the previous round of financing ends on September 12, allowing fundraising to commence as early as September 13. The company has not specified a definitive financing plan, but market expectations suggest fundraising may occur during the mid-September window.
Q: What are the release date, performance metrics, and iteration pace of Tencent's Hunyuan 4 Preview model?
A: Hunyuan 4 Preview was released in late August, earlier than originally anticipated. It features significant expansions in parameter count and context length, with internal blind tests showing performance slightly superior to GLM-5.3 and Kimi K3. The iteration speed has accelerated: Hunyuan 3 Preview was released in April, followed by the official version in July, indicating improved efficiency in model R&D.
Q: What is the strategic positioning and key focus of ecosystem building for Tencent's WorkBody open platform?
A: WorkBody is positioned as an open Agent ecosystem connecting hardware, applications, and developers. It supports continuous access for multiple hardware devices; on the application side, it allows industry partners to build AI workbenches based on the platform; developers can contribute skills and participate in distribution. The company is accelerating the development of payment and distribution capabilities to drive its evolution from an office assistant to an Agent operation layer.
Q: What are the investment implications of Tencent's AI strategy, and what are the key subsequent observation points?
A: The Hunyuan 4 Preview has marginally alleviated market concerns regarding model capabilities and iteration speed; WorkBody and Hunyuan form a data flywheel. If Hunyuan becomes a State-of-the-Art (SOTA) model, it can evolve into a Model-as-a-Service offering; if not SOTA, the WorkBody ecosystem still retains commercial value. Key areas to monitor subsequently include Hunyuan model performance, WorkBody progress, WeChat AI, as well as new product and strategic announcements at the Global Digital Ecosystem Conference in October.
Q: What were the highlights of Meituan's Q2 2026 performance, and what is the business guidance for Q3?
A: Q2 results exceeded expectations, primarily driven by the gross margin of the in-store business recovering to 30%; this improvement stemmed from Douyin competition capturing lower average-ticket orders, and the company proactively deferring some marketing expenses to Q3. Q3 guidance is relatively weak: food delivery unit economics (UE) are expected to decline quarter-on-quarter to 0.1, with in-store gross margin guided at 25%, mainly due to seasonal factors, increased rider subsidies, and no significant easing of the competitive environment.
Q: What are analysts' views on Meituan's stock price,support leveland their assessment of recent catalysts?
A: Analysts have adjusted the target price from 120 to 110, believing the stock price will oscillate within the 70–90 yuan range, with 70 yuan serving as a key support level. Recent catalysts to watch include the Bank-Futures Conference on September 22, where updates on Qwen 4.0 and related CAPEX guidance are expected.
Q: Regarding the performance of the Extra and F5.1 models, the competitive landscape of domestic models, and the expansion of paid scenarios, how should we assess the risks of price wars and user willingness to pay?
A: Evaluations of Extra and F5.1 are ongoing, making it difficult to conclude whether they fall short of expectations at this stage. Model competition will become stratified: companies with SOTA capabilities can efficiently distill high cost-performance versions, whereas companies with only cost-effective models will struggle to break through technical barriers and face pressure on gross margins, rendering pure price wars unsustainable. Following the correction of market mismatches, AI penetration rates will rise. Beyond AI coding, user willingness to pay is clear in high-knowledge-density industries such as finance, law, and healthcare. Given the higher difficulty of monetization on the consumer side, model companies are continuously shifting their strategic focus toward productivity scenarios.
Q: What are the core factors currently most concerning to buy-side investors regarding AI companies?
A: Buy-side investors prioritize model capability levels, followed by Annual Recurring Revenue (ARR) growth and gross margin levels; as long as model capabilities gain market recognition, valuation tolerance remains high. Short-term stock price fluctuations are primarily influenced by expectations regarding the pace of financing.