Skip to content

NVIDIA Puts Its Vera CPU to Work Designing Its Next Chips

NVIDIA is deploying its Vera CPU across the chip design workflows behind its next CPUs and GPUs, with Cadence and Synopsys tools running up to 1.5x faster in early tests.

NVIDIA superchip board upright, two exposed processor dies flanked by memory modules, on black
Credit: NVIDIA

NVIDIA has put its Vera CPU to work designing its own next generation of CPUs and GPUs, running the verification and simulation software that sets the pace of any chip project.

The company is tuning electronic design automation tools alongside Cadence and Synopsys, and two of those tools carry the headline number. Cadence Jasper, a formal verification platform, and Synopsys VCS, which simulates how a design behaves long before anything reaches a fab, each ran up to 1.5x faster on selected workloads in NVIDIA's own testing, with VCS given the same core count on both sides of the comparison.

What NVIDIA does not say is what Vera was 1.5x faster than. No competing processor is named, the workloads are described only as selected and production-class, and the testing is called early. The figure establishes that tuned builds of Jasper and VCS run well on Olympus cores. It does not yet establish that they beat the x86 machines filling most EDA compute farms.

The workload choice is deliberate. Logic simulation and formal verification lean on fast individual cores and low memory latency rather than parallel throughput, which is why they have resisted the GPU acceleration that reshaped other parts of the design flow. NVIDIA concedes the point, noting that several critical EDA workloads "remain heavily dependent on CPU performance." Verification is also the stage that eats the calendar: engineers spend years running regressions and hunting corner cases before a design is ready for tapeout, so time saved there compounds through everything downstream.

Vera pairs 88 Olympus cores with an LPDDR5X memory subsystem and a coherent on-die fabric carrying 164 MB of unified L3 cache across every core. The less obvious detail is topology. Each socket presents as a single NUMA domain on a monolithic die, where chiplet-based x86 servers can expose several per socket. Engineers running those machines spend real effort pinning threads and memory to avoid remote accesses, and a flat topology removes that tuning problem instead of optimizing it.

The larger consequence sits with the tool vendors rather than with NVIDIA. Once Cadence and Synopsys have tuned verification software for a server CPU that is not x86, that porting work is done for anyone who wants to run it, not only for the customer who asked. NVIDIA says it will carry the approach into Rosa, its next CPU, built on a core it calls Rigel, which leaves Vera's first serious job as designing its own replacement. Whether the 1.5x survives outside NVIDIA's regression farms is a question only Cadence's and Synopsys's other customers can settle.

Share this story

Thomas Albrecht

Thomas Albrecht covers processors and graphics silicon for techshooked, from desktop CPUs to the accelerators behind modern AI. His standard is numbers-first: document the test conditions, report the results a spec sheet leaves out, and judge a chip on measured performance rather than launch slides.