Nvidia Corporation Unveils Vera CPU and Rubin GPU for AI
Nvidia Corporation introduced the Vera CPU and Rubin GPU, forming a new rack-scale system designed to accelerate agentic AI and reinforcement learning workloads.
Nvidia Corporation has introduced a new hardware ecosystem centered on the Vera CPU and Rubin GPU, specifically engineered to accelerate agentic AI and reinforcement learning. The Vera CPU utilizes the new Olympus microarchitecture and an Arm-based design to minimize agent loop latency, reducing the time AI agents spend transitioning between GPU inference and CPU-driven orchestration. Featuring 88 cores and 176 threads on a monolithic die, the processor supports LPDDR5X memory and CXL 3.1, with Nvidia Corporation claiming it delivers twice the performance of current x86 processors in agentic sandbox tasks.
Complementing the CPU is the Rubin AI GPU, manufactured by TSMC using a 3nm process. The Rubin chip contains 336 billion transistors and employs HBM4 memory to reach a peak bandwidth of 22 TB/s, representing a 2.8x increase over the previous Blackwell architecture. It includes a third-generation Transformer Engine and enhanced Tensor Cores to optimize Mixture-of-Experts models and reduce token generation latency.
These components are integrated into the Vera Rubin NVL72 rack-scale system, which combines 72 Rubin GPUs and 36 Vera CPUs. Deliveries for these systems are expected to begin in late 2026. The platform has already secured ecosystem support from OpenAI and Perplexity, while Los Alamos National Laboratory reported achieving up to 7X faster performance in scientific computing workloads compared to Intel Sapphire Rapids-based systems.