Agent era CPU roadmap battle: NVIDIA bets on "faster single core," AMD bets on "more concurrency."
What Nvidia and AMD are competing for is not only whose CPU is faster, but who gets to define the benchmarking standards for CPUs in the AI era.
Nvidia recently disclosed the most detailed technical information about its Vera CPU architecture. This processor is equipped with 88 self-developed Olympus ARM architecture cores, memory bandwidth of 1.2TB/s, and on-chip interconnect bandwidth of 3.4TB/s. Nvidia’s core argument is: the workflow of AI Agents involves repeated CPU-GPU interactions—tool calls, code execution, data retrieval, task orchestration—each step depends on the completion of the previous one, therefore the single-core speed and latency of the CPU directly determine the response efficiency of the entire Agent.
On July 23rd, according to Chasing Wind Trading Desk, BofA Securities analyst Vivek Arya pointed out in his latest report that the release of Vera introduces a new evaluation framework for the industry: “maximum single-threaded performance at scale”. This stands in direct contrast to AMD’s longstanding chiplet multi-core stacking approach. The bank characterizes the debate as: “Is AI for agents bottlenecked by single-task completion time, or by how many concurrent tasks a single rack can run?” AMD will host its AI 2026 Tech Day this Thursday, which will mark its first formal public response to this framework.
The backdrop to this debate is that the server CPU market is being reshaped by AI demand. It is expected that by 2030, the market size will expand about fourfold from current levels, reaching $170 billion. This is not a zero-sum competition, but an emerging incremental market. Whoever establishes industry-recognized performance standards first will take control of pricing and the narrative.
Nvidia’s logic: Agents running fast is more important than how many can run simultaneously
Nvidia describes Agent workflows concretely: when an Agent completes a task, a large number of serial interactions occur between the CPU and GPU. Each call must wait until the previous step is completed before proceeding. This means delays at any stage will accumulate and magnify, ultimately slowing down the output efficiency of the entire AI factory.
Under this logic, single-core performance is not just a number on a spec sheet, but a key variable that directly affects GPU utilization—the faster the CPU, the less time the GPU spends waiting, and the higher the overall system resource utilization. In other words, the faster the single core, the faster the response speed of the entire chain.
Another design choice of Vera worth noting: Nvidia chose a monolithic die rather than a chiplet-separated architecture, arguing that the former provides better “scalable consistency.” This stands opposed to AMD’s long-term commitment to chiplet architecture, marking a deep divergence in the two companies’ underlying architectural philosophies.
Furthermore, Vera is not an isolated product, but a part of Nvidia's “co-design” AI infrastructure ecosystem, working together with Rubin GPU, Groq LPX, Spectrum switches, BlueField storage/network cards. This system-level integration is Nvidia’s true moat beyond simple hardware comparisons.

AMD’s rebuttal: In real AI production environments, the competition is about concurrent capacity
AMD’s position is built on a different definition of “real production environments.”
From AMD’s perspective, a production-grade AI system is not a single Agent advancing tasks serially, but more like a distributed software platform comprising databases, APIs, vector storage, orchestration engines, caching, and middleware. In this scenario, the bottleneck is not the speed of completing a single task, but how many workflows can be handled simultaneously within a fixed power budget.
AMD provided a set of concrete calculations: In a simulated 100-kilowatt rack deployment scenario, the rack-level throughput of EPYC 9965 (Turin) is about 2.4 times that of Nvidia Vera; next-gen EPYC 6 (Venice) is expected to reach 3.3 times.
AMD’s logic is: higher throughput density means more users served and more requests handled with the same energy consumption—this is the true cost function for cloud AI deployment.
AMD x86 or Nvidia ARM: Software ecosystem is the hidden battlefield
The CPU architecture debate also extends to instruction set level.
Nvidia’s choice of ARM architecture implies: as long as microarchitecture is excellent enough, instruction set compatibility can become less important. However, AMD and Intel will likely continue to emphasize another aspect after Thursday—AI is expanding from model inference to enterprise software workflows, and the enterprise software stack (databases, middleware, security platforms, business applications) has accumulated decades of optimization, verification, and compatibility in the x86 ecosystem.
This isn’t simply a technical debate. As AI workloads increasingly integrate into existing enterprise IT systems, the historical accumulation of the software ecosystem may be harder to replace than peak hardware performance. Whether Nvidia’s Vera can penetrate these scenarios largely depends on the maturation speed of the ARM software ecosystem.
Battle over standards: Who will define the next CPU KPI?
Two frameworks, two KPIs — essentially, it's two companies competing for control of the industry narrative.
Nvidia’s framework centers around latency, single-threaded progress, GPU utilization. AMD’s framework centers around concurrency, throughput, service density.
The key question: Which metrics will the market eventually use to purchase server CPUs?
BofA’s view is that AMD’s event this Thursday is less about unveiling benchmarks to overwhelm competitors and more about persuading the industry to accept its evaluation system. Whichever framework data center buyers adopt will gain pricing power in this $170 billion market.
Nvidia maintains a buy rating for NVDA, target price $350 (current price $207.29).
~~~~~~~~~~~~~~~~~~~~~~~~
The above content is from Chasing Wind Trading Desk.
For more detailed insights, including real-time analysis and frontline research, please join [Chasing Wind Trading Desk Annual Membership]
Risk Disclaimer & Exemption ClauseThe market involves risks, investment requires caution. This article does not constitute personal investment advice and does not take into account the individual user’s specific investment goals, financial situation, or needs. Users should consider whether any opinions, viewpoints, or conclusions in this article are relevant to their particular situation. Investing based on this is at your own risk.