share_log

Moonshot AI fully open-sources Kimi K3! It topped the trending chart of open-source communities within 30 minutes, prompting cloud providers both domestically and internationally to collectively announce Day 0 support.

cls.cn ·  Jul 28 11:27

① Moonshot AI released the model weights, technical report, and key infrastructure technologies underpinning the training of the Kimi K3 model;

② Kimi K3 topped the Hugging Face trending chart within half an hour;

③ Institutions noted that trillion-parameter-scale large models are driving the industry’s deployment architecture toward ultra-node clusters, creating new opportunities for domestic AI chips.

Moonshot AI has open-sourced its most capable model to date.

At 11 p.m. on July 27, Moonshot AI released the model weights and technical report for Kimi K3, and open-sourced the critical infrastructure technologies supporting Kimi K3 training: MoonEP, FlashKDA, and AgentEnv. This means anyone can now download and deploy the Kimi K3 model—freely using it for internal R&D or embedding it into end-user products.

Kimi K3 was launched in the early hours of July 17. It is a 2.8-trillion-parameter mixture-of-experts (MoE) model with native vision understanding capabilities and support for a one-million-token context window.

It is also the world’s first open-source model at the 3-trillion-parameter scale, specifically designed for advanced AI applications such as long-context programming, knowledge work, and complex reasoning.

Following its open-source release, Hugging Face CEO Clem Delangue posted that Kimi K3 topped the platform’s trending chart within 30 minutes, amassing over 4,000 likes—‘the fastest launch growth rate ever!!’

On the same day Kimi K3 was open-sourced, multiple companies across the AI ecosystem announced integration with Kimi K3.

Domestically, Huawei’s Ascend CANN announced that the full Ascend 950 series and Atlas A3 products support Kimi K3 deployment, noting that Kimi K3’s 896-expert EP deployment scenario fully leverages the advantages of Ascend’s ultra-node architecture. Additionally, Qujing Tech completed Day-0 inference adaptation of Kimi K3 on the Ascend Atlas A3 using the open-source inference engine SGLang.

According to Alibaba Cloud, its Zhenwu M890 ultra-node instance has achieved Day-0 compatibility with Kimi K3. The two parties will further deepen collaboration on domestic computing power, and both Qwen AI Platform and Alibaba Cloud Bailian will offer Kimi K3 via model APIs.

Internationally, AI infrastructure providers including NEBIUS, Baseten, and Fireworks also announced Day-0 support for Kimi K3. NEBIUS stated, ‘Kimi K3 is the first openly available model to reach state-of-the-art performance levels—a major milestone in the evolution of open models.’

Prominent AI programming company Cursor announced it has integrated Kimi K3; Cognition, the developer of digital employee Devin, stated that Kimi K3 is now available in the Devin desktop client and command-line interface (CLI), noting: 'On the Frontier Code1.1 benchmark, Kimi K3 is the first open-source model we have tested whose performance approaches frontier-level capabilities.'

Model weights, a 47-page technical report, and key infrastructure technologies are all publicly released.

Released alongside the Kimi K3 model weights are its technical report and model training methodology.

According to the announcement, the following technical details can be found in the technical report:

KDA+AttnRes: A 3:1 mixture of KDA and GatedMLA enables efficient long-context modeling, enhanced by block-level attention residuals to improve cross-layer information flow.

Stable LatentMoE: Each token activates 16 experts out of 896 routed experts, maintaining training stability under extreme sparsity through SiTU-GLU and Quantile Balancing.

MoonViT-V2: The vision encoder is trained from scratch using next-token prediction without contrastive pretraining, achieving baseline performance comparable to SigLIP initialization while enabling a more stable optimization process.

Post-training and evaluation: Large-scale synthetic tasks across three domains—general reasoning, general agents, and programming agents—supported by reinforcement learning infrastructure capable of handling million-token contexts, along with comprehensive evaluation results from nearly 20 internal benchmarks.

The capabilities of the model are underpinned by a robust training system—the infrastructure (Infra) layer. Moonshot AI provides detailed descriptions of three key Infra technologies supporting Kimi K3 training: MoonEP, FlashKDA, and AgentEnv. These cover critical components ranging from high-performance communication and operators to distributed reinforcement learning environments, which are essential for Kimi K3’s training efficiency and stability. FlashKDA was previously open-sourced, while MoonEP and AgentEnv are officially open-sourced with this release.

MoonEP: MoonEP is a high-performance communication library developed by Moonshot AI specifically for extremely large, fine-grained Mixture-of-Experts (MoE) models, enabling expert-parallel communication to achieve peak efficiency even under load imbalance.

FlashKDA: FlashKDA is a high-performance KimiDelta Attention kernel implemented by Moonshot AI. On NVIDIA H20 GPUs, it achieves a prefill speedup of 1.72–2.22× compared to the flash-linear-attention baseline and can serve as a direct drop-in replacement backend for flash-linear-attention.

AgentEnv: AgentEnv is a sandbox system co-developed by Moonshot AI and KVCache.ai for large-scale execution of agent environments. It provides a high-fidelity, strongly isolated sandbox for post-training of Kimi K3, with flexible support for rapid snapshotting, restoration, and forking—enabling efficient handling of massively parallel agent workflows and training tasks.

Open-source models enter the 3T era, continuously driving demand for cloud and computing infrastructure.

On July 24, Jensen Huang posted his first tweet on social platform X, sharing an open letter initiated by a16z and co-signed by NVIDIA along with several technology companies including Perplexity AI and Ollama. The letter explicitly advocates that 'cutting-edge open-source and closed-source models should coexist and develop in parallel,' highlighting the multiple values of open-source models—safety, innovation, and AI sovereignty—and asserting that the industry will simultaneously advance both cutting-edge closed-source and open-source models, thereby continuously boosting demand for cloud and computing infrastructure. Even with open-sourced base models, large model vendors retain advantages in low-level optimization, and accompanying open-source components are accelerating industry adoption.

Guojin Securities noted that domestically developed open-source models continue to achieve technical breakthroughs, with Kimi K3—the world’s first publicly available 3T-scale model—ushering in the 3T era for open-source large language models. The advent of 3T-scale models is driving a shift in industry deployment architectures toward super-node clusters, creating new opportunities for domestic AI chipmakers. Specifically, Kimi K3 officially requires a minimum of a 64-GPU cluster for efficient inference. At the 2026 World Artificial Intelligence Conference (WAIC), domestic chip vendors such as Huawei Ascend and MetaX showcased large-form-factor super-node products. Competition in the 3T model era now extends beyond single-card performance to encompass integrated system capabilities—including chips, high-speed interconnects, full-system clusters, and inference software stacks—meaning domestic chipmakers with comprehensive super-node offerings and strong ecosystem compatibility stand to benefit significantly.

Guosheng Securities stated that the capabilities of open-source models have significantly advanced, raising the industry-wide ‘baseline capability threshold.’ The firm remains optimistic about accelerated downstream application penetration and token consumption in the second half of the year. According to the firm, as open-source frontier models rapidly approach closed-source state-of-the-art (SOTA) performance, the ‘minimum capability baseline’ accessible across the entire industry has been substantially elevated, enabling downstream application developers to leverage near-frontier capabilities at lower costs. In terms of investment positioning, the firm favors:

(1) Application/Agent chains with real-world use cases, data moats, and engineering capabilities;

(2) Model-layer platforms with proprietary ecosystems and commercialization flywheels (e.g., cloud services), such as Alibaba, Tencent, Zhipu AI, and Minimax;

(3) Open-source scaling driving increased training and inference demand, with domestic computing infrastructure providers benefiting in tandem.

Editor/melody

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.