According to reports, NVIDIA is intensifying development of Nemotron 4, an open-source model with at least one trillion parameters, aiming to reduce its reliance on major clients such as OpenAI and cloud hyperscalers. The company has not only significantly increased its cloud compute capacity commitments to $28 billion through server sale-and-leaseback arrangements but also formed the 'Nemotron Alliance'—including partners like Mistral and Cursor—to co-build an ecosystem. On Tuesday, NVIDIA launched Nemotron 3.5 Lightning, a 30-billion-parameter model with 3 billion active parameters, alongside NeMo Switchyard, an open-source model routing library, targeting efficient agent execution and intelligent multi-model orchestration, respectively. Dell announced an expansion of its collaboration with NVIDIA, integrating these two newly announced NVIDIA technologies into its Deskside Agentic AI platform to extend agentic AI capabilities to on-premises enterprise deployments.
NVIDIA is accelerating its push into the competitive frontier of open-source AI models in an effort to expand demand for its GPUs—though this strategy simultaneously places it in a delicate position of direct competition with its own customers and investment targets.
According to a report by The Information on Tuesday, the 11th of this month, NVIDIA aims to bring the largest version of Nemotron 4 to performance parity with the world’s leading open-source models. Multiple employees revealed that the model is expected to have at least one trillion parameters—roughly double that of its current largest model, Nemotron 3 Ultra. In an email, Kari Briski, NVIDIA’s Vice President of Generative AI, stated, 'NVIDIA is investing in Nemotron because we believe every enterprise and every country needs access to state-of-the-art open-source models.'
NVIDIA’s current chip demand is highly concentrated among a handful of cutting-edge labs and cloud service providers, including OpenAI, Microsoft, and SpaceX—entities that are increasingly investing in their own AI chips. The advancement of Nemotron 4 aims to broaden NVIDIA’s customer base to a wider range of enterprise users, thereby reducing its dependence on these key clients. Anastasios Angelopoulos, CEO of AI model benchmarking firm Arena, remarked, 'No matter which company produces an excellent open-source model, NVIDIA wins.'
Following the news, NVIDIA’s stock rose nearly 2% in U.S. pre-market trading but opened lower after an initial gain of over 2%, turning negative by midday.
On the same day the news broke, NVIDIA officially announced Nemotron 3.5 Lightning, optimized for long-running AI agents, and unveiled NeMo Switchyard, an open-source model routing library, signaling the company’s accelerated expansion of the Nemotron model ecosystem. Dell also announced on the same day an expanded collaboration with NVIDIA, integrating both technologies into its Dell Deskside Agentic AI platform to enable agent deployment on enterprise local and desktop environments.
From trillion-parameter frontier models to efficient agent-focused models with 3 billion active parameters, and further to model routing and enterprise-local infrastructure, NVIDIA is building a comprehensive AI ecosystem spanning models, software, and deployment platforms. The announcement of Nemotron 4 development signals that NVIDIA is no longer content with being the world’s largest supplier of AI chips—it is now directly challenging top players in the open-source model space.
Targeting Trillion-Parameter Scale, Compute Investment Triples
According to multiple employees involved in the Nemotron project, NVIDIA plans for its largest Nemotron 4 model to have at least one trillion parameters—roughly double that of its current largest model, Nemotron 3 Ultra, released in June this year. Although this parameter count would still fall far short of leading Chinese open-source models, NVIDIA emphasizes that its compression technology enables smaller models to achieve superior performance.
NVIDIA’s resolve is evident in the substantial expansion of its financial commitments to computing power. The company acquires AI computing capacity by leasing back AI servers from cloud service providers that purchase its chips. As of April this year, the total value of NVIDIA’s multi-year cloud service commitments has risen to $28 billion by early 2031—approximately triple the amount disclosed a year earlier. Of this, around $7 billion is allocated for the current fiscal year (ending January 2028), which, although far below the investment scales of OpenAI and Anthropic, already exceeds the historical total funding raised by most leading open-source labs in the U.S. and China.
The Nemotron project is also continuously expanding in terms of research personnel. NVIDIA’s previous major model had a research paper co-authored by as many as 570 individuals, and multiple employees indicated that even more people will be involved in Nemotron 4. 'Everyone wants to get involved now,' said a former employee.
Forming the 'Nemotron Alliance' to onboard a broad range of open-source ecosystem partners
To accelerate the development of Nemotron 4, NVIDIA has established a collaborative network called the 'Nemotron Alliance,' comprising open-source developers such as Reflection, Cursor, Thinking Machines, and Mistral. While advancing their own model projects, these partners also contribute training data, evaluation support, and model design proposals to Nemotron 4.
Additionally, according to a source directly involved in the collaboration, AI model training startup Prime Intellect has contributed 300,000 simulated environments to the project to assist model training. AI coding tool startup Cognition is also part of the alliance; informed sources reveal that the company has held discussions with NVIDIA regarding the provision of code training data, though this collaboration has not yet been officially announced.
Alliance members are driven by diverse motivations. Some companies aim to expand their market share by leveraging NVIDIA's push for an open-source ecosystem; others hope to influence the development direction of a high-quality model they can use themselves at relatively low cost, thereby avoiding the burden of bearing training compute expenses on their own. Vincent Weisser, CEO of Prime Intellect, positions this alliance as a collective effort to counter the monopoly of a single 'god-tier' model, rather than a zero-sum race to claim the title of the best open-source model.
The boundary between customers and competitors is becoming increasingly blurred.
The advancement of Nemotron 4 has placed NVIDIA in a delicate structural tension. On one hand, OpenAI has long been a major driver of demand for NVIDIA chips, with NVIDIA having invested $30 billion in the company. On the other hand, Nemotron 4 is explicitly positioned as a lower-cost alternative for enterprises, directly competing with the commercial models of leading labs like OpenAI. Meanwhile, the alliance includes companies such as Reflection AI and Thinking Machines—open-source startups in which NVIDIA has already invested.
However, for NVIDIA, this competitive dynamic does not necessarily constitute a zero-sum game. Anastasios Angelopoulos, CEO of Arena—an AI model evaluation firm—stated: 'No matter which company produces an excellent open-source model, NVIDIA wins.' NVIDIA’s core rationale is that the more vibrant the open-source ecosystem and the more diverse its participants, the stronger the overall demand for GPUs.
Currently, the Nemotron series has achieved some market adoption. Palantir announced in June this year a partnership with NVIDIA to deploy Nemotron models for U.S. government clients. However, according to benchmark data from institutions such as Arena and Artificial Analysis, NVIDIA’s current largest model, Nemotron 3 Ultra, ranks only second among U.S. open-source models—trailing behind Thinking Machines’ newly released Inkling model—and falls outside the global top 40 overall.
The release timeline remains unclear.
Despite the project's scale, the release date for Nemotron 4 remains highly uncertain. Multiple employees indicated that NVIDIA has made several decisions regarding pretraining data and architecture but has not yet finalized specific specifications or a launch date, and the final training phase has not yet commenced, which is expected to take several months. Two employees believe the model could be released as early as late autumn this year, while others think it may come even later.
Meanwhile, NVIDIA this week released a smaller variant of the Nemotron series—Nemotron 3.5 Lightning—designed to run AI agent tasks with maximum speed and efficiency. NVIDIA also launched free model routing software to help enterprises more easily build custom model routing tools that automatically assign specific AI tasks to the optimal and most cost-effective model for processing.
Bryan Catanzaro, Vice President of Applied Deep Learning Research at NVIDIA, stated in a podcast this January that investing in Nemotron is 'critically important to our company’s future.' NVIDIA CEO Jensen Huang also publicly expressed support for open-source models in an email last month, stating that they 'promote safety, cybersecurity, scientific advancement, and national security.'
Nemotron 3.5 Leads the Way: 3 Billion Active Parameters Optimized for Long-Running Agents
On the same day news of Nemotron 4 emerged, NVIDIA officially launched Nemotron 3.5 Lightning.
Unlike Nemotron 4, which pursues 'larger' scale, Lightning prioritizes greater efficiency.
The model employs a Mixture-of-Experts (MoE) architecture with approximately 30 billion total parameters, though only about 3 billion parameters are activated per token. It is specifically designed for AI agents requiring long-duration operation and repeated model calls. NVIDIA claims that Nemotron 3.5 Lightning can achieve up to a 4x improvement in token generation speed.
Its role is not to replace the most powerful frontier models, but rather to serve as an 'execution-oriented' model within multi-agent systems.
For example, in a complex enterprise AI system, large models can handle task planning and complex reasoning, while Nemotron 3.5 Lightning can take on specialized tasks such as large-scale code review, security monitoring, data processing, and alert analysis.
This is also the value of the Mixture-of-Experts (MoE) architecture: the model has a large total parameter count, but only a small subset is activated during each inference, thereby reducing the per-token computation cost.
NVIDIA aims to address real-world challenges that arise when AI agents operate at scale—where a single agent may invoke a model not just once, but dozens or even hundreds of times consecutively.
Therefore, for AI agents, it is important that a model be “sufficiently intelligent,” but being “sufficiently inexpensive and fast” is equally critical.
However, a fourfold increase in token generation speed does not necessarily translate to a fourfold acceleration of an entire agent task.
Third-party benchmarks show that Nemotron 3.5 Lightning demonstrates a significant speed advantage in pure model generation; however, when integrated into a full agent workflow, the overall task speed improves by approximately 30%. This is because model inference is only one component of an agent system—tool calls, API response times, and task orchestration among multiple agents can all introduce new bottlenecks.
This also sets the stage for NVIDIA’s concurrent launch of NeMo Switchyard.
From “Model” to “Routing”: NVIDIA Begins Addressing Bottlenecks in Agent Systems
NeMo Switchyard is an open-source model routing library.
Its core idea is that developers can automatically route requests to the most suitable model based on the characteristics of each task—without needing to rewrite their existing applications.
For example, complex reasoning tasks can be assigned to large models, while high-frequency tasks such as code review, classification, and data extraction can be delegated to efficient models like Nemotron 3.5 Lightning.
The significance of this approach lies in the fact that enterprises no longer need to use the most expensive model for every agentic task.
For future multi-agent systems, what truly matters may not be 'which model is best,' but rather which model should perform which task.
By simultaneously launching Nemotron 3.5 Lightning and Switchyard, NVIDIA has effectively signaled its strategic view on agent infrastructure: as the number of models continues to grow, orchestration and coordination among models will form a new software layer.
This layer, in particular, can help NVIDIA expand its ecosystem influence.
Even if enterprises ultimately deploy models from other vendors, as long as those models are integrated into NVIDIA’s routing, inference, and deployment framework, NVIDIA can still capture value from the entire AI workflow.
Dell expands collaboration: Bringing Nemotron and Switchyard to the enterprise desktop
NVIDIA’s model strategy is further solidifying its presence in the enterprise segment.
On August 11, Dell announced an expanded partnership with NVIDIA, integrating Nemotron 3.5 Lightning and NeMo Switchyard into the Dell Deskside Agentic AI platform. Dell stated that this platform targets enterprise-grade agentic applications, emphasizing local execution, low latency, cost control, and data security.
This means NVIDIA’s move on that day was not merely an isolated model release.
On one end are cutting-edge open-source models such as Nemotron 4, targeting the trillion-parameter scale; on the other are smaller, highly efficient models like Nemotron 3.5 Lightning, designed for local and enterprise agent workloads.
Dell further integrates the latter with on-premises hardware and enterprise IT environments.
For enterprise customers, running agents locally is attractive because sensitive data does not need to frequently leave the internal corporate network, reliance on cloud-based APIs can be reduced, and latency and inference costs become more controllable.
Dell previously positioned Deskside Agentic AI as a gateway connecting enterprise desktops to data centers. Its solution is built on Dell AI Factory with NVIDIA and supports scaling from local agent development to data center–scale deployment.
Therefore, by incorporating Nemotron 3.5 Lightning and Switchyard into its platform, Dell is effectively reinforcing another NVIDIA strategy:
enabling agents to operate not only in the cloud but also directly within enterprise employees’ working environments.
If an increasing number of enterprises adopt a hybrid model combining local small models with cloud-based large models in the future, NVIDIA stands to benefit across multiple layers—including GPUs, models, routing software, and enterprise infrastructure.
Editor/rice