Skip to content

Nvidia begins shipments of Vera CPU as dedicated hardware for agents re-engineers task management and code execution

Share
Nvidia begins shipments of Vera CPU as dedicated hardware for agents re-engineers task management and code execution

Listen to this article

Read by Anchor

Nvidia has begun commercial shipments of its custom Vera CPU, marking the company's first central processor built specifically for AI agent workloads, moving past the traditional model that confined AI tasks to GPUs alone. Ian Buck, Nvidia's vice president of hyperscale and HPC, delivered the initial server units to major cloud providers and research labs, including AWS, Oracle Cloud Infrastructure (OCI), Anthropic, OpenAI, and SpaceX AI.

This hardware shift reflects the current phase of AI systems, where models no longer merely generate statistical text, but must execute Python code, handle tool calls, manage sandboxed environments, and coordinate retrieval across long contexts. These sequential and concurrent tasks architecturally fall on the CPU rather than the GPU, creating bottlenecks in traditional infrastructure that was designed simply to pack cores without considering real-time agent flow.

The Vera processor features 88 custom Olympus cores designed by Nvidia, with memory bandwidth reaching 1.2 TB/s, delivering up to 1.8 times higher single-core performance on agent workloads compared to standard designs. The processor connects to the Vera Rubin GPU via second-generation NVLink-C2C interconnects within the NVL72 system, using a unified memory architecture alongside BlueField-4 DPUs and Spectrum-X networking, doubling energy efficiency relative to previous architectures.

In cloud infrastructure, AWS expanded its partnership with Nvidia to include plans for two million additional GPUs alongside the Vera architecture, while Oracle announced plans to deploy hundreds of thousands of Vera processors starting in 2026, making it the first cloud provider to offer them at hyperscale for high-throughput reasoning workloads. Among research labs, SpaceX AI has begun evaluating the processor for reinforcement learning and agent-directed simulation pipelines, while infrastructure leads at OpenAI and Anthropic have focused on testing the chip's capacity to accelerate agent compute scaling.

This shift reshapes operating economicsFor technical teams and enterprises across the region, particularly in the Gulf and Egypt, a significant portion of current enterprise agent costs stems from latency in code execution and cloud container environments. The availability of cloud instances combining fast CPU processing and high memory bandwidth translates directly into lower latency for autonomous systems and higher throughput for branching tasks without bottlenecks, prompting regional engineers to recalibrate their cloud provisioning criteria and calculate token costs based on supporting CPU efficiency rather than GPU capacity alone.

Don't miss the next story

Subscribe for updates