NVIDIA Vice President Ian Buck: The Evolution of Computing Power in the Agentic AI Era Through the Vera Rubin Platform
Restructuring the Hardware Cost Reduction Logic in the AI Era
"The most effective way to reduce the cost of a word is not to lower the price of GPUs or use low-cost networks, but to deliver significantly improved computing performance with each generation. If the performance of a node or rack is increased by 10 to 30 times, this performance improvement will directly affect the numerator of the calculation formula, thereby reducing the cost of each word in the entire data center."
Analysis
Traditional IT procurement practices typically control costs by lowering hardware unit prices. However, in the era of AI factories, the improvement in computing power performance is exponential (10 to 30 times), and its impact on Total Cost of Ownership (TCO) far exceeds that of minor adjustments to chip unit prices. Therefore, improving performance to reduce the cost per unit of computing power is far more efficient than simply purchasing low-priced hardware.
Agentic AI eliminates human intervention, leading to a surge in computing power consumption.
"The intelligent agent has removed humans from this interaction loop, and the rate at which computing resources are consumed depends entirely on the speed at which the intelligent agent can change course and think. Previously, service designers could rely on the gap of 'human reading and typing takes time' to allocate computing power, but now the intelligent agent asks questions and changes course extremely quickly. In the AgentX benchmark test, the computing power requirement has increased by 100 times in pure numerical terms."
Analysis
The shift from Chat mode to Agentic AI has brought about a paradigm shift in computing power demand. In the past, the underlying assumptions of cloud service scheduling were based on the limitations of human reaction and input speed. Agents, however, remove the human-in-the-loop buffer, allowing computing power to enter a seamless and high-frequency full-load state.
The core metrics for evaluating AI chips have shifted.
“For each generation of computing power products, we consider not only absolute word performance, but also the total word throughput per megawatt (MW). Improving this metric can enhance the revenue-generating ability of the AI factory.”
Analysis
The evaluation system for AI hardware is being reshaped. Industry assessment standards have shifted from pursuing single-chip FLOPs or single-card benchmark scores to evaluating the power conversion efficiency (Tokens per Megawatt) of data centers. The core value of computing power is no longer limited to peak performance, but is reflected in the efficiency of converting electrical energy into actual revenue.
Discover 40% of the hidden capacity of data centers through software.
"By dynamically managing power consumption at the rack and firmware levels through software, ensuring that it never exceeds the threshold, up to 40% of racks can be deployed under the same reserved power envelope, directly increasing throughput to 1.2 billion words per second."
Analysis
The static deployment model of data centers, which reserves power based on the maximum peak power consumption of hardware, has been broken. Under the objective condition of limited power supply, by using software-defined power technologies such as MaxLPS, GPU deployment density can be increased by 40% within the existing power envelope without building new data centers or increasing power quotas.

Ian Buck: The Evolution of Computing Power in the Agentic AI Era: A Perspective from the Vera Rubin Platform
Table of contents
- Introduction: The Power of 25 Years of Computing Evolution and Data Transformation
- Paradigm Shift: From Chat Benchmarks in 2023 to Agent AI Workloads
- Understanding Vera Rubin: A Full-Stack Architecture for Agent Computation
- Extreme Latency and Throughput: Groq LPU Integration with Olympus Core
- Ecosystem and Expansion: NVLink Fusion, Intelligent Storage and MaxLPS Power Management
- Looking ahead: Delivering AI infrastructure roadmap on an annual schedule
Introduction: The Power of 25 Years of Computing Evolution and Data Transformation
Ian Buck: Today I'll be sharing how we're advancing infrastructure for the AI era of intelligent agents, which essentially represents a new era in the meaning of reasoning. I'll share some insights from my work perspective and introduce the innovative technologies we're developing, giving the community here an understanding of the products we're building. I'll also be detailing Vera Rubin, explaining the significant changes that have occurred in the past few years, and even just the last year, and the measures we've taken at the high level.
I believe the field of computing is incredibly vast. Drs. Ian and Patterson have explained how computing has transformed their perspective, extending from early magnetic core storage to today and the future. This truly demonstrates the importance of computing infrastructure and computing power: transforming data, acquiring information, performing calculations, outputting results, and then feeding back the transformations. This is the essence of value creation. Storing information is important, transmitting information is important, but value is generated in the computation and in the process of transforming actual data into computational data.
From our perspective, this platform is massive. Having worked at NVIDIA for nearly 26 years, it all clearly began with the launch of a completely new computing platform: CUDA. Since then, its applications have expanded dramatically. This graph illustrates the diversity of people using our GPUs, encompassing hyperscale cloud providers, AI-native companies, startups, major AI labs, enterprise customers, and edge computing. It blends AI models with traditional analog computing, or rather, a combination of both, as they now support each other.
Hugging Face now boasts over 3 million models, with approximately 3,000 models accounting for about 90% of downloads. The diversity of models continues to expand, with workloads extending from computer vision to physical sciences, life sciences, speech, natural language processing, and robotics. We possess AI models for quantum computing and simulation, as well as multimodal and cutting-edge models exploring entirely new architectures. Beyond these, we offer a comprehensive library of simulations and applications for computation, covering fields such as computational lithography, decision making, medical imaging, weather analysis, and genomics.
Computing has become an indispensable capability for scientific research, industrial production, consumer interaction, and enterprise computing, spanning AI and analog computing. We currently have over 8,000 major CUDA applications running on these platforms, and we have been building this platform for approximately 25 years.
Paradigm Shift: From Chat Benchmarks in 2023 to Agent AI Workloads
What's the current situation? This is Vera Rubin. It was designed to be the next-generation inference and agent platform. Obviously, we continue to support large-scale training, which is also a capability we can provide. But what makes all of this possible is the AI factory, and Vera Rubin is designed for that.
It's built on a rich ecosystem. We can track over 10 million developers worldwide developing on NVIDIA GPUs. This includes cloud providers, network service providers, OEMs, and ODMs. Over 300 major companies have built the supply chain to help deliver this platform—it's amazing to watch all of this happen.
It boasts extremely high versatility and substitutability. I believe this is one of the reasons why it can bring tremendous value and help globally: you provide a computing platform that can run all workloads and every model, whether it's an open-source model, a closed-source model, or a cutting-edge model. It can support pre-training, post-training, inference, agent AI, and accelerated computing, all of which are invoked simultaneously within agent workloads.
Furthermore, it is incredibly durable. David mentioned that the GPUs in GCP are still delivering the original Voltas, which dates back a long time, almost a decade ago. You can still see our GPUs, released a decade ago, creating value. In fact, if you track the hourly instance rental price of the A100 we released early in the COVID pandemic, you'll find that it has actually increased over time, and demand remains strong.

So what changes have occurred at the software level? The most significant change is in the workload. A large portion of the initial benchmarking was based on Chat, a groundbreaking workload in 2023 that later became the foundation for benchmarks like MLPerf Inference. It had approximately 1K input tokens, and the KV cache for system prompts and other user data was likely around 4K in size. People interacted with it in a multi-turn dialogue. You asked a question, it responded, and then you might ask follow-up questions, averaging about three rounds of dialogue. This was a human-in-the-loop model; there were no agents at the time, but it became the basis for benchmarking and understanding our value in how well such workloads performed, whether it was the DeepSeek model or a closed-source interface like ChatGPT.
Today, agent AI is entirely different, presenting challenges beyond imagination. Its requirements are extremely stringent, increasing by 100 times at the purely numerical level. This data comes from the AgentX benchmark; while it may not represent everyone, compared to our current benchmarks and previous inference tests, the average input size reached 142,000 words. Moreover, it's dynamic; you need to support short inputs ranging from 1K, 10K, 100K all the way up to 142K, 200K, and even longer. Each input length requires a completely new set of operator kernel optimizations, tuning, and mathematical adjustments, greatly increasing the complexity of software optimization.
The KV cache has become incredibly large. Models are starting to support up to 1 million terms. You might start small, but as agents engage in multi-turn interactions (now it's not a human in the loop, but the agent constantly asking itself questions), the context keeps accumulating. Now all of these KV caches need to be managed across the entire data center.
Finally, the number of interaction rounds has exploded. Previously, service designers could rely on the fact that humans needed time to read the answer before providing the next prompt, making task scheduling relatively easy. Now, the agent removes humans from this loop, and the rate at which computational resources are consumed depends entirely on the agent's self-reinvention speed, which is extremely fast.

In AgentX benchmarks, they only set up about four sub-agents, while in our commercial services we have tens of thousands of agents, all of which need to perform intelligent routing. The model now has to make decisions and invoke all these different sub-agents across the platform, and the KV cache needs to be retained at every stage.
Examining the workload's structure reveals its far greater complexity. What used to be simply lexical input, running the model, and lexical output now involves scheduling CPUs and managing dynamic Kubernetes containers via Helm Charts. We have an LLM handling context and observation data, another layer for inference, and yet another layer deciding which sub-agents or tools to invoke in the next round.
Finally, we need to run all these tools: some are CPU-based, some require GPUs, some are simulations, and some are accelerated computations. On top of that, we have CPU clusters managing and monitoring security and governance. We don't want agents to hallucinate, generate false ideas, or waste time and lexical revenue, as that would cause the system to crash. All of this needs to be managed within memory context windows, which has led to an explosion in the size of these context windows, which serve as the model's working knowledge base.
Understanding Vera Rubin: A Full-Stack Architecture for Agent Computation
This was the original design philosophy behind Vera Rubin. When we released it, we weren't talking about launching a single chip or reducing the cost of a single LLM model; we were talking about being able to run entire workloads.


On the base platform side, Vera Rubin NVL72 builds upon and expands upon what we did for Grace Blackwell. We also added an acceleration package based on Groq LPU technology when we wanted to achieve faster speeds from a user lexical perspective. For these high-value workloads, we can further enhance platform performance.
Above all, the CPU is crucial for the intelligent agent. We must be able to complete all tool calls and calculations quickly, efficiently, and at full capacity, so as to return the results to the user rapidly. We absolutely do not want the CPU to slow down the entire technology stack.
Finally, all KV caches need to be managed in intelligent storage. As various contexts flow within the system, all data needs to be managed, governed, and integrated into the software-defined network, while indexing and vector retrieval of the KV cache are performed in the background for optimization. This is why we launched BlueField for storage, partnering with all our storage partners who build these systems, providing them with SmartNICs and compute platforms.
Now let's look at large-scale networks. These models are so large that a single GPU simply can't handle them. In fact, everyone is currently working on decoupled inference. We discussed Dynamo software last year, and it's now widely used. Everyone is separating the prefill and decode phases and managing all the key-value caches on top of them to achieve optimal terminator rate and throughput. Connecting all these components requires the support of network infrastructure. Of course, all these facilities can also be used for RL and post-training phases.
These are the results of running the newly released SemiAnalysis agent benchmark on Vera Rubin. Through full-stack co-design with the entire ecosystem across both open-source and closed-source models, it achieves a 30x increase in AI factory throughput for agent workloads. In fact, depending on where you observe the Pareto curve, the performance improvement can even reach up to 60x in some scenarios.

On the left is the total word throughput per megawatt (MW). All of these data centers are limited by megawatt (MW) power consumption. For each generation of products, we consider not only absolute word performance, but also the key metric of total word throughput per megawatt (MW). Improving this metric enhances the revenue-generating ability of the AI factory.
If you're running a model like DeepSeek v4, you'd want a thinking rate of 100 morphemes per second or higher to compute answers and return quickly. In fact, you might want to go further to 200 to 250 morphemes per second, allowing the agent to think more deeply and provide answers faster across 65 rounds of interaction, without making the user wait too long.
Furthermore, we must reduce the cost per million units. This graph compares the cost of generating units using Grace Blackwell and Vera Rubin. As you can see, the cost drops significantly due to performance improvements. The most effective way to reduce costs is to deliver stronger performance with each generation. If I simply lower the price of GPU chips or use a cheaper network, I'm only improving a small part of this massive data center. If I increase the performance of that node or rack by 10 or 30 times, this performance improvement directly impacts the numerator of the equation, thereby reducing the cost per unit across the entire data center.
Extreme Latency and Throughput: Groq LPU Integration with Olympus Core
What if we want to think faster, more proactively, and execute rapid interactions? In many application scenarios, intelligent agents are crucial, and we want to place them on the critical path for real-time business analytics or cutting-edge computing. We can further enhance the performance of Vera Rubin. From Blackwell to Vera Rubin, we've improved the overall user experience. But if we want to push the limits, from hundreds of terms per second to thousands of terms per second per user, we can leverage Groq technology to further unlock performance.



These are the latest test results. For models with 100K long contexts, such as Gemma 4 or Qwen 3.8, we can achieve 3400 words per user per second on Gemma 4 and 2529 words per user per second on Qwen 3.8. This speed is remarkable in scenarios with large contexts. The longer the context, the more the model needs to retrieve the entire dialogue history to generate each new word, which increases the computational load, but this is necessary for intelligent agent applications.
Tool invocation is extremely critical. Globally, there's a shortage not only of data centers and GPUs, but also of CPUs. While previously Grace CPUs connected to Blackwell were used for tool invocation, we are now seeing a need for efficient, fast, dedicated CPUs for agent computing. We extracted CPUs that were originally connected to Rubin GPUs and built a dedicated Vera CPU rack.
What is Vera? It uses NVIDIA's custom-designed Olympus core. Its design goal is to achieve high speed, outputting answers as quickly as possible while each core is running at full load. We achieved this through a custom Olympus core.
To ensure all cores operate at high speed without bottlenecks, extremely high bandwidth is maintained between cores. This is a large, single-chip compute die, allowing each core to access memory without multi-hop transfers across chiplets, resulting in extremely low memory access latency. Combined with an LPDDR memory subsystem, it not only boasts low power consumption but also 40% lower latency than other solutions on the market. Memory latency often plays a dominant role in CPU computation.
This differs from the CPU behavior before the cloud era. Past CPUs boasted high single-threaded performance but had fewer cores, resulting in lower overall throughput. In the cloud computing era before AI, the metric became the cost per core, driving up the number of cores while single-threaded performance became secondary. This explains the surge in core count, as cloud service customers primarily inquired about the price per core for running web hosting and databases.

Vera is designed to maintain high throughput across 188 cores (176 cores with SMT enabled) while maintaining high single-threaded performance and ensuring that lexical latency is not affected.
These results were validated using Signal65 with Stanford University's Terminal Bench benchmark. We ran Terminal Bench CPU-side tasks across more than 80 tests, and Vera demonstrated a 1.6x performance improvement. The agent completed tool calls 60% faster, with up to a 2x improvement in processing speed on core agent tasks such as compiling code, running Git, and code verification.
Customers have begun seeking to deploy this product. Perplexity Space discovered that task execution was often limited by the speed of sandbox creation. By introducing Vera, they observed an overall speed improvement of 1.5x in sandbox startup, execution of rapid computations, and shutdown.
ClickHouse is a real-time analytical database used by leading labs for security and governance. They ran its public benchmark suite on Vera and ranked it number one.
Vera has now been officially released, with early adopters including Oracle Cloud Infrastructure (OCI) and several system vendors.
Ecosystem and Expansion: NVLink Fusion, Intelligent Storage and MaxLPS Power Management
NVLink is a key element for enabling high-speed AI and inference. From no NVLink to 8-way interconnect, and then to the full NVLink 72, it has significantly improved the Pareto performance curve. NVIDIA is a networking technology company, and we are opening up NVLink to everyone through the NVLink Fusion program. Chip designers only need to design their own XPU accelerator chips, and by partnering with NVIDIA, they can interface with our NVLink Chiplet IP and NVHBM technology, thus directly integrating into the complete NVLink ecosystem.

Early partners include the startup The Matrix, which is using NVLink Fusion for its next-generation Raptor XPU to connect Vera and NVLink Switch racks, building a complete NVLink 72 architecture incorporating The Matrix chips. Amazon AWS and Annapurna Labs are also collaborating with us to integrate custom IP into NVHBM substrate dies, freeing up 30% to 40% of compute resources on their custom chips while leveraging NVIDIA SerDes for higher bandwidth and lower latency.
At the data center level, as we focus on improving inference efficiency per megawatt (MW) or gigawatt (GW), a core question is how many GPUs can be accommodated within a data center. The power consumption characteristics of workloads are dynamic. Even on the DeepSeek-R1, GPU power consumption fluctuates between 600W and 1900W depending on the stage, resulting in underutilization of reserved power resources.

To address this issue, we collaborated with the ecosystem to develop software called MaxLPS (Max Land Power Shell). MaxLPS dynamically manages and limits power consumption at the Vera, Rubin, and NVSwitch levels, as well as at the rack and firmware levels. By working in conjunction with data center management software, operators can deploy up to 40% more racks within the same power envelope (e.g., scaling from 4,000 racks to 5,600 racks in a gigawatt-scale data center), increasing token throughput to 1.2 billion tokens per second.
Lambda validated this software stack. Dave Ward pointed out that at the same megawatt (MW) capacity, the deployable GPU density was significantly improved, with the B200 seeing an increase of about 25% and the Vera Rubin seeing an increase of up to 40% in computing resources.

Looking ahead: Delivering AI infrastructure roadmap on an annual schedule
NVIDIA consistently adheres to a strict one-year roadmap. Following Vera Rubin, our next presentation will introduce Vera Rubin Ultra, followed by the Feynman architecture and Feynman Ultra. You can be assured that NVIDIA will maintain its annual release schedule, continuously advancing the development of AI lexical factories, inference, and data center-scale agent infrastructure.
Thank you all for your time.
This article is sourced from Andy730.
Risk warning and disclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.