share_log

Ahead of AMD's AI conference, NVIDIA showcased the achievements of its CPU 'Vera Rubin,' now in full-scale production, with over 300 partners deploying it.

wallstreetcn ·  Jul 22 01:34

NVIDIA announced at the end of May that Vera had entered full production, stating that the chip can complete specific tasks 1.8 times faster than traditional x86 CPUs. On Tuesday this week, NVIDIA emphasized that, rather than simply stacking more cores, Vera prioritizes single-threaded performance, inter-core communication bandwidth, and memory access latency. The company also stated that the Vera Rubin NVL72 is now ramping up in mass production globally, with partners including CoreWeave, Google, and Microsoft already beginning to deploy related racks. Benchmark tests conducted by CoreWeave show that the Vera Rubin NVL72 delivers a 10-fold increase in tokens per megawatt compared to the Blackwell architecture.

NVIDIA is extending its competitive battle from the GPU market into CPUs.

On Tuesday, June 21 (U.S. Eastern Time), just ahead of AMD’s AI-related event, NVIDIA unveiled significant progress on its next-generation Vera Rubin platform: the Vera Rubin NVL72 has entered volume production ramp-up. NVIDIA stated that the platform’s supply chain spans over 350 factories across 30 countries, with more than 300 partners already involved.

Simultaneously, NVIDIA released performance data for its Vera CPU and the Vera Rubin platform in AI agent workloads, aiming to demonstrate that as AI evolves from merely 'answering questions' to autonomously planning, invoking tools, and executing tasks, CPUs are becoming a new battleground in AI infrastructure.

The full specifications, benchmark results, and architectural details of NVIDIA’s data center CPU product, Vera, constitute critical information required by potential customers for comprehensive chip evaluation. The company stated that Vera chips were delivered to customers including OpenAI, Anthropic, and SpaceX in June.

The launch of Vera represents the latest step in NVIDIA’s vertical integration strategy. The company aims to sell complete rack-scale computing systems featuring its proprietary chips rather than individual components alone. According to Wolfe Research estimates, the average unit price of Vera chips is approximately $5,000, with projected shipments reaching around 1.3 million units this year. For AMD and Intel, NVIDIA’s move directly threatens their core interests in the server CPU market.

This is not the first time NVIDIA has announced that the Vera CPU has entered mass production. At the end of May this year, NVIDIA declared that Vera had already reached full-scale production, claiming the chip delivers 1.8x the performance of traditional x86 CPUs on specific tasks. In mid-May, NVIDIA had already shipped the first Vera CPU systems to customers such as Anthropic, OpenAI, and SpaceX AI. The current announcement focuses further on the production ramp-up of the entire Vera Rubin platform, customer deployments, and real-world performance in production environments.

Rise of AI Agents Reignites CPU as the 'Critical Bottleneck'

NVIDIA’s core rationale for betting on CPUs lies in the evolving nature of AI applications.

Traditional generative AI primarily focused on generating responses based on user inputs, with GPUs handling large-scale model computations while CPUs managed relatively peripheral scheduling tasks. However, as AI agents rapidly advance, AI systems now autonomously decompose tasks, invoke external tools, execute code, access data, and iteratively evaluate outcomes.

In this process, CPUs must frequently handle numerous low-latency, real-time tasks. NVIDIA contends that AI agents do not represent workloads solely dependent on GPUs: each agent’s runtime environment, tool invocation, task orchestration, and long-context data retrieval all require active CPU involvement.

When NVIDIA previously launched Vera, it noted that agent-based AI is creating a new 'CPU moment.' Its assessment is that as AI systems shift from 'answering questions' to 'taking actions,' the role of CPUs within the broader AI infrastructure will significantly increase.

NVIDIA even projects that the long-term market size for server CPUs could reach $200 billion. While this figure appears highly aggressive compared to the traditional server CPU market, NVIDIA’s rationale is that the future market boundary for CPUs may no longer be confined to conventional enterprise servers but could expand into AI inference, AI agents, reinforcement learning, data processing, and various control and orchestration tasks within AI infrastructure.

This is also why NVIDIA is attempting to redefine the competitive rules for CPUs.

Rather than competing on core count, NVIDIA is betting on 'single-core speed.'

Vera’s design approach differs markedly from that of traditional server CPUs.

NVIDIA states that Vera is its first CPU specifically designed for agent-based AI, featuring 88 NVIDIA-designed Olympus cores and 1.2 TB/s of memory bandwidth. NVIDIA previously indicated that Vera delivers a 50% improvement in single-core performance over conventional CPUs; in tests released at the end of May, Vera achieved 1.8 times the task completion speed of x86 CPUs in specific workloads.

In its latest disclosures, NVIDIA further emphasized Vera’s core design philosophy: rather than simply stacking more cores, Vera prioritizes single-threaded performance, inter-core communication bandwidth, and memory access latency.

NVIDIA claims that Vera’s custom Olympus cores deliver 2x single-threaded performance, 3x higher inter-core bandwidth, and 40% lower memory latency compared to competitive chiplet-based designs. The goal is to enable AI agents to complete tasks faster and return computing resources to GPUs more quickly, thereby improving overall utilization across the AI infrastructure.

The underlying business logic is straightforward: GPUs are among the most expensive and scarce computing resources in AI data centers. If CPUs process tasks too slowly, GPUs may remain idle. NVIDIA aims to minimize this 'idle' time by optimizing CPUs specifically for AI workloads.

In production environment tests conducted by DeepInfra, NVIDIA reported that Vera can support up to 1.6 times more concurrent AI agents and achieve up to 2.2 times faster task orchestration speeds. It should be noted that these figures come from customer-specific test environments and do not imply that Vera outperforms all existing CPUs across general-purpose workloads.

The competitive focus has shifted from a single chip to an entire 'AI factory.'

The Vera CPU is only one part of NVIDIA's broader systems strategy.

Within the Vera Rubin platform, NVIDIA integrates the Vera CPU with the Rubin GPU, Groq 3 LPX, Spectrum-6 networking chips, ConnectX-9 SuperNIC, and BlueField-4 into a co-designed system.

NVIDIA stated that the Vera Rubin NVL72 is now ramping up volume production globally, with partners including CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure already deploying related racks. NVIDIA also noted that its supply chain spans more than 350 factories across 30 countries and involves over 300 partner companies.

In customer testing, CoreWeave reported that after running DeepSeek-R1 tests on the Vera Rubin NVL72, token generation per megawatt-second improved tenfold compared to the previous-generation Grace Blackwell NVL72. Google Cloud has launched A5X instances based on the Vera Rubin NVL72, which NVIDIA says can achieve lower inference costs and higher token throughput per megawatt in specific scenarios.

This means NVIDIA’s competition is no longer limited to individual accelerator chips such as AMD’s Instinct GPUs or Intel’s Gaudi processors.

NVIDIA aims to control the entire infrastructure stack: CPUs handle orchestration, GPUs perform computation, networking chips enable high-speed interconnectivity, and software manages scheduling and optimization—ultimately delivering integrated systems at the rack or even data center level to customers.

For AMD and Intel, the threat of this competitive approach lies in the fact that NVIDIA may not need to engage in fully symmetric competition in the traditional CPU market. Instead, NVIDIA targets the fastest-growing segments of AI servers, initially capturing workloads such as agentic AI, reinforcement learning, and high-performance AI inference, and gradually expanding the application scope of its CPUs.

AMD and Intel’s traditional strengths now face new challenges.

Intel and AMD have long held dominant positions in the server CPU market.

However, NVIDIA now believes that AI infrastructure is reshaping the value hierarchy of CPUs. In the past, competition among server CPUs centered primarily on core count, general-purpose computing capability, and overall throughput; in the era of agentic AI, low latency, single-thread performance, memory bandwidth, and task orchestration efficiency are becoming increasingly important.

This is precisely the market gap NVIDIA sees as an opportunity to enter.

In terms of product form factor, Vera can function either as a standalone CPU or be paired with NVIDIA GPUs to form the Vera Rubin system. NVIDIA also positions it as a critical component linking AI agents to GPU compute resources.

NVIDIA has already listed customers including AI companies such as Anthropic, OpenAI, and SpaceX AI, as well as cloud service providers like ByteDance, CoreWeave, and Oracle Cloud Infrastructure. Server manufacturers such as Dell, Hewlett Packard Enterprise, Lenovo, and Supermicro are also developing systems based on Vera.

This gives NVIDIA an advantage distinct from traditional CPU vendors: it can offer customers comprehensive solutions combining GPUs, CPUs, networking, and software.

However, this does not mean NVIDIA can easily displace AMD and Intel.

Competition in the CPU market is not solely about chip performance—it also encompasses software ecosystems, customer certifications, server compatibility, and long-term supply relationships. Large cloud service providers and enterprise clients, in particular, typically require several years to transition to a new hardware platform.

Therefore, whether Vera can truly expand beyond AI-specific use cases into the broader server market will ultimately depend on whether customers can achieve sufficiently compelling performance and cost advantages.

NVIDIA is opening a new revenue curve beyond GPUs.

NVIDIA’s concentrated release of Vera Rubin’s production readiness and performance details just before AMD’s AI event also carries clear competitive intent.

Over the past few years, NVIDIA has nearly monopolized the most lucrative segment of AI infrastructure—GPU-based computing. However, as AMD continues to expand its AI accelerator portfolio and Intel seeks a foothold in the AI chip market, investors are increasingly questioning whether NVIDIA can sustain its high growth trajectory.

Meanwhile, AMD’s and Intel’s CPU businesses have also regained market attention due to new demand expectations driven by agentic AI.

NVIDIA’s response is this: if the value of AI data centers is expanding from standalone GPU computing to full-fledged AI factories, then NVIDIA need not cede the CPU market to its competitors.

From this perspective, Vera is not merely another server CPU launched by NVIDIA, but rather a further extension of its vertical integration strategy.

NVIDIA aims for customers to purchase not individual GPUs, CPUs, or networking chips, but an integrated infrastructure stack capable of running AI models and AI agents out of the box.

If agentic AI indeed becomes the primary driver of the next wave of computing demand, NVIDIA may end up competing not just for GPU market share, but for dominance across the entire AI server computing value chain.

This also means that AMD and Intel may soon find themselves competing not only against each other, but against an NVIDIA that already possesses capabilities spanning GPUs, CPUs, networking, software, and full-system solutions.

Editor/Stephen

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.