- Nvidia’s Olympus core architecture prioritizes single-threaded IPC over frequency, using a 10-wide decoding interface, a deeply out-of-order midcore, and a graphics prefetcher tuned for workloads with many agent AI branches and pointers.
- Vera swaps chiplet-style core density for a monolithic 88-core array and a single-NUMA-per-socket design, a deliberate move toward agentic AI workloads that Nvidia’s own engineers admit comes at the expense of performance for legacy workloads.
- Nvidia’s self-reported SPEC CPU 2026 results show a per-core advantage figure of between 70% and 80% compared to AMD’s EPYC 9755 server CPU.
Nvidia’s Vera CPU has a lot to prove, as it represents the company’s first real attempt at the AI server CPU market, where traditional vendors Intel and AMD, along with third-party Arm-based vendors, are seeking a slice of an increasingly lucrative data center pie.
To that end, Nvidia has released its most detailed white paper yet on the chip, promising substantial performance improvements over the competition. Vera is the company’s first server processor built around an all-internal core design, a departure from the original Arm cores that powered its Grace predecessor.
As Vera approaches general availability in the second half of 2026, Nvidia is making its case on one specific front: not raw core count, but sustained performance per core under load, the metric it claims is most important for agent AI.
Latest videos ofTechnologyRadar
What’s really under the hood of Nvidia’s Vera CPU?
At the core of each Vera CPU is Olympus, Nvidia’s first custom server core, built to the Armv9.2 instruction set but designed in-house rather than being derived from Arm’s original Neoverse designs, as was its predecessor, Grace.
Nvidia’s second-generation data center CPU is a completely redesigned design built around a single goal: leading agent workload performance.
Rather than follow Intel and AMD down the chiplet route, Nvidia packages all 88 cores (176 threads) onto a single monolithic chip, with a dual-socket configuration offering 176 cores and 352 threads in a single system.
Nvidia has extensively detailed the core design and several publications have since delved into the microarchitecture. The interface runs a 10-wide decoding engine combined with a neural branch predictor that can resolve up to two taken branches per cycle, designed to handle the large instruction footprints and irregular control flow of interpreters, compilers, and agent runtimes.
The middle core combines an extensive allocation and renaming engine with a large reordering buffer and dependency breaking techniques including memory renaming and value prediction.
The execution engine dynamically schedules integer, vector, floating-point, cryptographic, load, and storage resources, while the cache subsystem adds a graph prefetcher targeting pointer lookup access patterns that defeats conventional streaming prefetchers.
That design has not gone unanswered. AMD has responded with estimated numbers for its 256-core Zen 6 “Venice” part, claiming a 3.3x rack-level advantage over Vera, although those numbers are extrapolated rather than measured.
That AMD responds first is not surprising. Within x86, it continues to gain ground on Intel, reaching 33.2% of x86 server CPU shipments in the first quarter of 2026 according to Mercury Research, up from 27.2% a year earlier.
With Arm reporting that its architecture now accounts for about 50% of CPU processing among major hyperscalers, competition is building from several directions at once. Vera could well set a new standard for agent AI.
Follow TechRadar on Google News and add us as a preferred source to receive news, reviews and opinions from our experts in your feeds.




