Full Breakdown
Nvidia’s Vera CPU and Rubin Platform Enter Mass Production, Targeting AI Agent Workloads
7/22/2026, 7:52:58 AM
Core Event
Nvidia announced that its Vera Rubin NVL72 racks—combining the Vera CPU and Rubin GPU—are in full production and deployed at several cloud providers and AI labs as of July 21, 2026. The same day the company released new architectural details and benchmark data for the Vera CPU, a custom Arm-based processor codenamed Olympus.
Background & Context
Nvidia has dominated AI hardware with GPUs. To capture more of the AI data-center value chain, it developed its own CPUs, first Grace and now Vera, aimed at “agentic” AI workloads that interleave GPU inference with CPU-driven orchestration.
Architecture & Specifications
- Core design: Monolithic die with 88 Olympus cores, each supporting spatial multithreading (72 threads per core).
- Front-end: 10-wide decode, neural branch predictor, 64 KB L1 instruction cache, 96 KB L1 data cache.
- Memory: LPDDR5X delivering up to 1.2 TB/s bandwidth; 40 % lower loaded latency versus chiplet designs.
- Interconnect: Second-generation Scalable Coherent Fabric with 3× core-to-core bandwidth over chiplet architectures.
- Rack composition: Each NVL72 rack houses 72 Rubin GPUs, 36 Vera CPUs, 18 compute trays, and 9 NVLink switch trays. The sixth-generation NVLink spine delivers up to 3.6 TB/s per GPU and 260 TB/s scale-up bandwidth per rack.
- Cooling: Liquid “dry cooling” eliminates fans and reduces water usage.
Performance Benchmarks
- Nvidia’s internal SPEC CPU 2026 run shows the Vera CPU scoring 3 % higher than AMD’s 128-core EPYC Turin (925 vs. 898) despite a thread disadvantage, with per-core IPC gains of up to 1.9× and branch-prediction improvements of 2.3×.
- Independent coverage notes the SPEC results are “estimated” and not directly comparable to publicly submitted scores.
- Customer tests: CoreWeave measured a 10× performance-per-watt improvement for the full Vera Rubin platform on the DeepSeek-R1 workload; the Vera CPU delivered 1.8× speedup on agentic SPEC workloads versus EPYC Turin. Perplexity reported a 1.5× faster sandbox job completion, and Prime Intellect observed a 30 % increase in RL sandbox throughput.
Customer Deployments & Use Cases
- As of July 21, 2026, Vera Rubin NVL72 racks are operating at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Bristol Myers Squibb’s DGX SuperPOD.
- Bristol Myers’ chief research officer Robert Plenge said the new system will let the company evaluate “dozens” of drug candidates instead of “10,” potentially cutting early-stage development time by up to 50 %.
Official Statements & Responses
- Buck framed the Vera platform as a way to maximize “intelligence per dollar” for post-training workloads.
Conflicting Reports & Gaps
- Nvidia’s SPEC CPU figures are presented as normalized per-core scores, while Tom’s Hardware points out that the underlying dual-socket EPYC baseline is not disclosed, making direct comparison difficult.
- The claim of “2× performance” lacks a clear reference point, and independent third-party validation of many benchmark claims remains pending.
Verbatim Quotes
- “When you host these things, you have to pay an electric bill,” — Greg Meyers, chief digital and technology officer
- “They're now producing three times the output or effectively nine trillion dollars of productivity,” — Ian Buck, general manager of Nvidia's hyperscale and HPC computing business
What’s Next
Nvidia plans to ship Vera Rubin racks throughout the second half of 2026 and has outlined a roadmap that includes larger configurations with up to 256 GPUs. A July 17, 2026 blog frames future AI workloads as “post-training” cycles, suggesting subsequent hardware generations will focus on further improving token-per-watt efficiency.
