NVIDIA's Vera CPU trounces x86 chips, such as AMD's EPYC, in Agentic AI workloads, delivering a 6x performance gain with 40% lower latency.
NVIDIA Vera's Monolithic CPU Design & Architectural Changes Achieve Over 2x Performance Gain Over x86 CPUs Such As AMD EPYC With Just 88 Cores, Beating 128 Cores
Today, NVIDIA is going all out by sharing its full performance disclosure of the Vera CPU against x86 offerings. We just talked about the architecture in detail over here, and now, we will look at the performance capabilities that this chip has on offer. NVIDIA is evaluating Vera across four key categories that matter the most in the Agentic AI era:
- Memory Bandwidth and Latency – Shows how Vera keeps cores supplied with data through high-bandwidth memory, low-latency access, and a high-throughput coherent fabric.
- Application Workload Performance – Demonstrates performance across agentic benchmarks, data analytics, graph processing, and other CPU-intensive workloads.
- Core IPC – Measures the impact of the Olympus microarchitecture, including the wide Front End, advanced branch prediction, deep out-of-order execution, and expanded execution resources.
- RL and Agentic AI Performance – Highlights Vera’s value for latency-sensitive, branch-heavy, and highly concurrent agent environments.
But before starting with the benchmarks, we want to state that NVIDIA has heavily used the AMD EPYC Turin "Zen 5" chip in its comparison, highlighting the monolithic & NUMA nature of its chip & how it significantly elevates its position over x86 chips.
Architectural Benchmarks
NVIDIA has divided the benchmarks into two categories, one that focuses on the architectural side of things and the others at how the architectures improve AI workflows. As such, the first architectural benchmark showcases Vera's IPC uplifts across various workloads that stress large instruction footprints, dense control flow, compiler and runtime behavior, and long dependency chains. Here, Vera offers up to a 1.9x uplift over AMD's Zen 5 (EPYC Turin), showcasing its strong single-threaded gains, a key driver for Agentic AI and RL apps.
The Branch Prediction on Vera's Olympus cores is up to 2.3x faster and 2x faster on average versus AMD's Zen 5. The Taken-branches per cycle see a 3.5x uplift.
Olympus is designed to sustain high taken-branch processing across dense and varied branch patterns, made possible by the neural branch predictor in addition to other special-purpose predictors. Olympus BPU (Branch Prediction Unit) operates at a higher effective rate while still maintaining a branch MPKI (Missed Prediction per1000 Instructions) advantage, delivering more useful instruction streams into the wide decode and execution pipeline. The result is better front-end utilization, fewer wasted cycles, and more useful instructions, achieving up to 3.5x higher taken-branches per cycle.
The Olympus core features a 64 KB instruction cache with a high-bandwidth fetch of 128 bytes per cycle out and a 10-wide decode without relying on an x86-style uOP cache, which offers up to a 2.4x gain over Zen 5 in Instruction Fetch Ops per cycle.
And lastly, we have the Backend Ops per cycle, which are up to 4.3x faster and over 3x faster on average versus Zen 5.
Olympus is designed to sustain high taken-branch processing across dense and varied branch patterns, made possible by the neural branch predictor in addition to other special-purpose predictors. Olympus BPU (Branch Prediction Unit) operates at a higher effective rate while still maintaining a branch MPKI (Missed Prediction per1000 Instructions) advantage, delivering more useful instruction streams into the wide decode and execution pipeline. The result is better front-end utilization, fewer wasted cycles, and more useful instructions, achieving up to 3.5x higher taken-branches per cycle.
AI Workload Benchmarks
The first benchmark is centered around memory latency. Here, NVIDIA states that traditional x86 CPUs these days rely on chiplet architecture. The chiplet design leads to shortcomings in the memory subsystems when many requests are made across cores. These cause bottlenecks to appear in the memory subsystem and the coherency fabric, and as a result, additional hops increase latency and reduce effective bandwidth, leading to performance drawbacks.
In the comparison, the AMD EPYC Turin 9755 CPU can be seen jumping from similar latency in the beginning (~120ns) up to 350ns as soon as memory bandwidth demand increases. Vera, on the other hand, can deliver 3x the bandwidth with 40x lower peak latency at a 90% memory utilization rate.
NVIDIA's Vera CPUs also offer much higher memory bandwidth per core. The Turin chip has 128 cores, with each core being fed roughly 3.1 GB/s of bandwidth. Vera has 88 cores, and each core is fed with around 12.7 GB/s of bandwidth, 4x higher than EPYC, and the per-core memory bandwidth is achieved up to 14 GB/s in certain use cases. Such per-core memory bandwidth characteristics are important for Agentic AI workloads where many threads operate on large, irregular data sets and can quickly become limited by memory throughput rather than compute resources.
Further impacting the performance of chiplet-based solutions is the addition of dozens of separate I/O and compute chiplets. When data moves from one core to another, it may need to leave the source core die, traverse the package fabric through an I/O die, reach a different core die, and then follow a similar path back for coherency responses. Each die hop adds latency, consumes fabric bandwidth, and creates variability depending on where the communicating cores are physically located. The result is higher core-to-core latency & impacts performance.
In the core-to-core latency heat map, you can see that NVIDIA Vera delivers 50% lower latency than the AMD EPYC 9755 chip. What's even more impressive is that Vera shows consistent latency across all of its core pairs, further emphasizing how its monolithic architecture is made from the ground up for AI workflows.
Next up, we move to Agentic benchmarks, which are the core reason why Vera was made in the first place. In Spec CPU 2006, NVIDIA measures the per-core performance for a fully loaded socket system. The software stack here measures performance across Python execution, code compilation, static analysis, and tool-driven software workflows.
Fully loaded per-core performance captures how well each core sustains throughput while sharing socket-level power, memory bandwidth, cache, and fabric resources.
The NVIDIA Vera CPU delivered up to 1.8x higher performance in Python versus AMD's EPYC. The rest of the benchmarks show up to 1.7x performance; that's almost 2x the performance bump vs x86 competitors.
In Graph Traversal algorithms, NVIDIA's Vera CPU offers a 2.6x performance uplift over AMD Turin. Graph workflows are bandwidth and cache-sensitive, and here, NVIDIA's memory and cache subsystems inside Vera are showcasing displaying their full potential.
NVIDIA is also showing off its Graph Traversal Core scaling in Page Rank workloads, which also stresses the core pipeline, fabric, and memory bandwidth. Here, Vera achieves a monumental 29.3x gain at a core count of 32.
Lastly, we have ClickHouse, which is a high-performance columnar database used for online analytical processing. Vera achieves a 20% uplift over the AMD EPYC Turin chip and shows it to be a great fit for scan-heavy, aggregation-heavy, and concurrency-sensitive analytics and ETL (Extract-Transform-Load) workloads.
NVIDIA's Vera CPU represents a significant architectural leap forward for high-performance computing, particularly in the demanding realm of Agentic AI workloads. By embracing a monolithic design rather than the chiplet-based approach common in modern x86 CPUs, Vera overcomes critical bottlenecks in memory subsystems, coherency fabric, and inter-core communication. This results in dramatically lower and more consistent latency, while delivering substantially higher memory bandwidth per core.
Gunning A $200b TAM Oppurtunity
The NVIDIA Vera CPU is aiming to become a highlight of the Agentic AI CPU race. The company has already said that Vera is in a position to become the leading CPU supplied in 2026, and there are plenty of early adopters that have started using Vera, such as OpenAI, Anthropic, SpaceX, & Perplexity. The CPU is being deployed across various supercomputer centers and cloud data centers, and will see full-fledged support from countless ecosystem partners.
Optimized from the ground up for irregular, bandwidth-hungry, and concurrency-intensive AI workflows, Vera not only sustains higher per-core throughput under full load but also sets a new standard for efficient, scalable processing in AI datacenters.
Follow Wccftech on Google to get more of our news coverage in your feeds.
