NVIDIA Vera CPU, the company's first processor aimed at AI agents, has moved past one-off demos and is now leaving the factory in quantity as an 88-core custom Olympus design, per NVIDIA.
The latest recipient is Amazon Web Services. NVIDIA's Ian Buck, who runs hyperscale and HPC, took an NVIDIA Vera CPU server and a Vera Rubin GPU to Seattle and handed them to Amazon EC2 vice president Willem Visser, with EC2 hardware-engineering vice president Supreeth Sheshadri on site. That visit sits on top of a freshly widened 16-year AWS partnership that NVIDIA says covers two million extra GPUs plus work to stand up Vera-based racks inside AWS.
AWS is not the first stop. Buck had already walked the same class of box into Oracle Cloud Infrastructure and into three AI labs (Anthropic, OpenAI, and SpaceXAI). Karan Batta of OCI said the cloud intends to field hundreds of thousands of NVIDIA Vera CPU chips as agent demand grows, and NVIDIA calls OCI the earliest provider putting Vera in at hyperscale. James Bradbury, who leads compute at Anthropic, described Vera as a useful piece of the agent-workload stack. SpaceXAI is still weighing it for reinforcement-learning jobs and agent-driven simulation pipelines.
On paper the custom Olympus design brings 1.2 TB/s of memory bandwidth and up to 1.8x faster per-core results on agentic AI jobs, per NVIDIA. The argument is practical: agents chew through sandboxes, tool calls, orchestration, and long-context retrieval, and that work lands on CPUs even when the model itself runs on GPUs. NVIDIA Vera CPU is also the host chip in Vera Rubin NVL72, tied over second-generation NVLink-C2C to a pair of Rubin GPUs that share a unified memory map.
Buck's line is that agentic AI is opening a new CPU chapter inside the AI factory as models shift from answering questions to taking actions, and that Vera is the part meant to keep that loop moving.













