Vera Rubin’s measured performance is revealed for the first time, with a significant boost! Nvidia “crashes the party” ahead of AMD’s conference.
```
On the eve of rival AMD's annual product launch, NVIDIA delivered a heavy blow, intensively unveiling real-world performance data of its next-generation Vera Rubin platform and officially revealing the full specifications of its self-developed Vera CPU chip.
On Tuesday, NVIDIA disclosed that the Vera CPU's performance on agent-style AI tasks is nearly double that of x86 chips, with latency improved sixfold. Meanwhile, early production tests from cloud computing partner CoreWeave show that for the DeepSeek R1 model, the token throughput per megawatt of compute power on the Vera Rubin NVL72 platform is ten times higher than the previous generation GB200 NVL72 system based on the Blackwell architecture.
The timing of these data releases is highly meaningful. AMD will hold its annual “Advancing AI” product launch event in San Francisco this Thursday, and NVIDIA’s concentrated performance data release is widely interpreted as deliberate market hype. For cloud vendors and enterprise customers evaluating next-generation AI infrastructure investments, this data batch will directly influence their purchasing decisions.
At the same time, NVIDIA also officially announced that the Vera CPU completed its first batch of deliveries in June, with customers including OpenAI, Anthropic, and SpaceX. This marks NVIDIA’s formal entry into the CPU market, directly challenging the traditional turf of AMD and Intel.
Vera CPU Debuts: NVIDIA Officially Enters the Server CPU Market
On Tuesday, NVIDIA unveiled the full specifications, benchmark results, and architecture information of its datacenter CPU product Vera—critical for potential customers’ comprehensive evaluation of the chip. The company stated the Vera chip has been delivered to clients such as OpenAI, Anthropic, and SpaceX as of June.
Vera is NVIDIA’s first server CPU designed autonomously at the core level, utilizing a custom microarchitecture codenamed “Olympus core” instead of ARM’s off-the-shelf designs.
Hannah Coutand, NVIDIA’s product marketing lead for Vera, explained the chip’s design focuses on single-core speed, high memory bandwidth, and low latency, aiming to “enable agents to return to the GPU as quickly as possible and keep the GPU always highly utilized.” On the hardware specs side, Vera’s power consumption is between 250 and 450 watts, with each chip supporting up to 1.5TB of low-power memory.
In the latest benchmarks, NVIDIA demonstrated a 1.9x performance increase for agent-style AI tasks over x86 chips, and a six-fold reduction in latency. NVIDIA also said Vera outperforms AMD’s flagship EPYC Turin CPU by nearly 100% in some industry-standard benchmarks. Previously, NVIDIA claimed Vera’s overall performance on AI agent tasks is 50% higher than x86 chips.
Vera’s launch is the latest step in NVIDIA’s vertical integration strategy. Wolfe Research estimates Vera chips average $5,000 apiece, with shipment volume expected to reach about 1.3 million units this year. Ian Buck, NVIDIA’s VP of hyperscale computing, stated agent-style AI makes CPUs “more indispensable,” predicting the server CPU market could ultimately reach $200 billion.
Vera Rubin Real-World Testing: 10x Energy Efficiency Leap
CoreWeave completed the industry’s first deployment and verification of Vera Rubin NVL72 in early June, covering power, cooling, networking, and compute full-stack confirmation. The data disclosed this time is the first public real-world performance result of the Vera Rubin NVL72 silicon.
Using the DeepSeek R1 model as a benchmark, with the same interactive response target, Vera Rubin NVL72 generates tokens per megawatt per second at ten times the rate of GB200 NVL72. CoreWeave also noted optimizations for Vera Rubin can be retroactively applied to GB200 NVL72 systems, boosting megawatt throughput by over four times in three months. NVIDIA says these achievements were validated with over 250,000 unique configurations and more than 1.4 million GPU hours of testing.
Architecturally, Vera Rubin NVL72 racks integrate 72 Rubin GPUs and 36 Vera CPUs linked via 260 TB/s fully-connected NVLink 6 and natively support NVFP4 precision. CoreWeave says this efficiency boost means customers can process more inference traffic within the same power budget or run equivalent loads at lower power, reducing per-token costs.
CoreWeave also revealed specific application scenarios: A global cybersecurity company expects to run threat detection inference with 10x per-watt performance; an autonomous coding agent company expects to scale agent tasks at much lower token cost; an AI-native search engine expects to expand real-time search services to more users without breaching response time limits.
Synergized Infrastructure Optimization: Integrated Hardware and Software Unlock Greater Performance
NVIDIA emphasized that the performance gains aren't solely dependent on the chip itself, but are the result of comprehensive hardware and software co-design strategies. By dynamically optimizing the full infrastructure and energy stack, NVIDIA claims it can deploy up to 40% more GPUs within the same power envelope.
For cooling, NVIDIA uses a 45°C closed-loop liquid cooling system, saving about 4 million gallons of water per megawatt annually compared to standard methods. Network-wise, sixth-generation NVLink 6 interconnect architecture delivers 2.3x higher simulated decoding throughput for large language models compared to Ethernet; the Spectrum-X platform achieves 1.6x faster remote direct memory access bandwidth, reduces switch quantities by 1.7x, ups optical power efficiency fivefold, and enhances reliability tenfold.
NVIDIA's latest Spectrum-6 platform has begun deliveries to AI factory customers like CoreWeave, Microsoft, Nebius B.V., SpaceXAI Corp., and Tesla. NVIDIA is currently ramping up shipments to Google Cloud, Microsoft Azure, Meta, Oracle Cloud Infrastructure, Dell Technologies, OpenAI, and CoreWeave customers and partners.
NVIDIA Remains a Challenger in the CPU Field
Despite Vera’s impressive performance numbers, NVIDIA still faces significant market share challenges in CPUs. Gartner analyst Kevin Knox says AMD is currently the main competitor in enterprise AI server CPUs. Reports indicate Intel holds about 66.8% market share in server CPUs, AMD about 33%, and AMD is steadily gaining share and building strong collaborations with hyperscale cloud providers.
Boosted by expected CPU demand from agent-style AI, AMD and Intel stocks have climbed 128% and 149% respectively this year, far outpacing NVIDIA’s roughly 8% increase, making them the standout chip stocks for 2026.
For now, NVIDIA’s main cloud service partner list only includes Oracle and hasn’t yet reached other mainstream providers. Hannah Coutand admits Vera is still in the “early adoption stage,” but OpenAI plans to begin large-scale deployment of Vera chips as early as this quarter.
Karl Freund, founder of Cambrian AI Research, summarized NVIDIA’s strategy: “Their goal is to free customers from reliance on Intel or AMD CPUs—and they’re committed to monetizing that revenue segment. They’re focusing on delivering a unique CPU that no one else currently offers in the market.”
Risk Disclosures and DisclaimerThe market carries risks; invest cautiously. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, viewpoints, or conclusions in this article suit their specific circumstances. Investment decisions based on this are at your own risk. ```