AMD fires back at Nvidia, claiming 256-core Zen 6 'Venice' CPU beats Vera by 3.3x in rack-level performance — company shares first estimated EPYC Venice benchmarks
Mind the footnotes on this one.
AMD has shared the first official benchmarks for its forthcoming EPYC 'Venice' CPUs, which will be the first chips to use the Zen 6 architecture. The flagship 256-core model hasn't been detailed in full, but AMD claims it offers 3.3 times the performance of the Nvidia Vera CPU in a rack-scale implementation with a fixed power budget of 100kW.
The constraints of this test completely change the framing of the results, so it's worth highlighting them first. AMD is looking at performance from the level of a rack, not an individual component on a single socket. AMD's results are modeled around a 100kW deployment, showing the performance across the rack rather than the performance on a single socket, or even a dual-socket system.
AMD did not, however, actually test all of these deployments. There's a lot of modeling going on here, which AMD details in its methodology paper that was published alongside the results. First, AMD estimated power based on the processor TDP and additional components, and it used that to calculate the number of nodes (2P system for each node) within a 100kW power budget. Then, it multiplied that number of nodes by single-node performance measured in a handful of benchmarks.
That's just the start of the stipulations. AMD doesn't have its hands on Vera, so the performance here is an estimate. AMD took benchmarks it had for Nvidia's Grace chip and multiplied them by a scaling factor of 1.63x based on Vera results published by Phoronix. AMD also says that its 256-core EPYC Venice results are derived from an estimated 1.7x scaling factor over the EPYC 9965, along with "internal testing."
AMD's methodology ends with this line: "[These results] are intended to provide directional comparison rather than direct measured rack benchmarks." As you can tell from the three paragraphs of stipulations above, that sentence carries a lot of weight on its shoulders. You can't simply scale the performance of a single "node" (a dual-socket system, in this case) up in a linear fashion. Interconnects, as well as thermal and power limitations, will become a factor as you scale up.
Regardless, AMD frames the results here around agentic AI, though the benchmarks it's using are focused on general-purpose data center tasks. The topline result is from SPEC CPU 2017, specifically looking at integer throughput. AMD used the following benchmarks, as well:
- Server-side Java based on SPECjbb 2015
- WRK Tool for load on an NGINX web server
- Redis-benchmark for in-memory workloads
- Memory caching with Memcached
- Database performance with TPROC-C on MySQL
These "results" are mostly a way for AMD to fire back at Nvidia after Phoronix published a list of results for Vera that were curated by Nvidia. AMD is laying the groundwork for its Advancing AI event next month, where we expect to hear more about Venice, Zen 6 more broadly, and AMD's enterprise roadmap in much greater detail.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.

Jake Roach is the Senior CPU Analyst at Tom’s Hardware, writing reviews, news, and features about the latest consumer and workstation processors.
-
bit_user Reply
Yes, mind the footnotes on all of them, Vera included!The article said:Mind the footnotes on this one.
This is like 20x what a free-standing electric oven uses, just to put that in perspective.The article said:in a rack-scale implementation with a fixed power budget of 100kW.
But the scaling factors they used should account for that stuff, to the extent that the Vera testing was realistic. The big unknown is how Vera scales to dual-CPU configurations, because Phoronix' testing was only on a single Vera CPU (even though the pictured card had two).The article said:You can't simply scale the performance of a single "node" (a dual-socket system, in this case) up in a linear fashion. Interconnects, as well as thermal and power limitations, will become a factor as you scale up.
IMO, all of these benchmarks and estimates are just rough notions. We'll just have to wait a few months until systems based on Vera and Venice actually get out into the wild, before we get a clear picture of how they really compare.
My own sense is that Vera might have faster per-core performance, but nowhere near enough to win on either the basis of performance or efficiency, per socket. However, NVLink might mean that Vera scales up better. That could enable it to win on performance density. -
DS426 Reply
I'm thinking the same. What would make Vera more interesting is if nVidia made a quad-socket design like Intel has available in some Xeon configurations while AMD remains at a dual-socket max. Ignoring memory-constrained scenarios, this could help green level the playing field in some workloads.bit_user said:...
My own sense is that Vera might have faster per-core performance, but nowhere near enough to win on either the basis of performance or efficiency, per socket. However, NVLink might mean that Vera scales up better. That could enable it to win on performance density. -
bit_user Reply
They did better than that!DS426 said:What would make Vera more interesting is if nVidia made a quad-socket design like Intel has available in some Xeon configurations
If you look at the connectivity, sure the chip-to-chip bandwidth between a pair of them is 1.8 TB/s (I think 900 GB/s per direction), but then you can link those 2P nodes in a larger coherent fabric (NVLink is cache coherent).
I didn't find an answer to what the inter-node bandwidth would be, but I did find this quote about Vera CPU-only systems:
"it will also be possible to get whole racks of Vera server CPUs in the Oberon racks with the ETL spine. (Meta Platforms is going to be an early customer for this.) If you do the math, that is eight Vera CPUs (possibly four two-way Vera-Vera nodes) in each sled, with 32 sleds in a Vera ETL racks. That is 256 CPUs with a total of 22,528 cores and 512 TB of main memory and 300 TB/sec of bandwidth across that memory. "
Source: https://www.nextplatform.com/compute/2026/03/19/driving-down-the-ai-system-roadmap-with-nvidia/5210195
So, that assumes 8 CPUs per sled, which is the maximum scaling that Xeon traditionally supports.
Well, let's look at that. AMD's Venice will have CPUs with 256 Zen 6c cores and 512 threads, but if they can only equip those CPUs with 4 TB of DRAM, that works out to 8 GB per thread (2x threads per core). Vera supports up to 1.5 TB of memory per CPU, which works out to 8.7 GB per thread (2x threads per core). So, the per-core or per-thread capacity is roughly the same.DS426 said:Ignoring memory-constrained scenarios, this could help green level the playing field in some workloads.
Next, let's look at density. If Nvidia can pack 8 Vera's + memory in the same space as two AMD Venice CPUs + memory, then it works out to 704 cores + 12 TB vs. 512 cores + 8 TB. So, the density argument goes in Vera's favor. -
thestryker Reply
I've wondered how much of the software they allowed for testing benefited from shifting to Spacial MT from Simultaneous MT. Just another question that can't really be answered until these systems are actually deployed and tested though.bit_user said:Yes, mind the footnotes on all of them, Vera included! -
bit_user Reply
I'm still suspicious that "Spatial MT" isn't just marketing spin, trying to cast a weakness as a strength.thestryker said:I've wondered how much of the software they allowed for testing benefited from shifting to Spacial MT from Simultaneous MT.
This tidbit is the best info I have yet to find about it:
https://www.realworldtech.com/forum/?threadid=227058&curpostid=227650 -
thestryker Reply
I hope we get to see these designs because I'd love to see what nvidia is sacrificing to get 8 CPUs per sled. If the way nvidia is doing it is indeed 4 of the double Vera CPU boards I don't see anything stopping an AMD vendor from doing similarly. The memory footprint is same width wise and RDIMMs are ~55mm longer. It looks like the Vera package is around 25-30mm narrower than the specs of the SP7 socket. Neither one of which is enough to prevent a 2x 2 CPU AMD design from being possible in the same amount of space.bit_user said:Next, let's look at density. If Nvidia can pack 8 Vera's + memory in the same space as two AMD Venice CPUs + memory, then it works out to 704 cores + 12 TB vs. 512 cores + 8 TB. So, the density argument goes in Vera's favor. -
thisisaname Reply
So not so much accurate benchmarks but more like a wet finger to judge wind speed and direction.bit_user said:Yes, mind the footnotes on all of them, Vera included!
This is like 20x what a free-standing electric oven uses, just to put that in perspective.
But the scaling factors they used should account for that stuff, to the extent that the Vera testing was realistic. The big unknown is how Vera scales to dual-CPU configurations, because Phoronix' testing was only on a single Vera CPU (even though the pictured card had two).
IMO, all of these benchmarks and estimates are just rough notions. We'll just have to wait a few months until systems based on Vera and Venice actually get out into the wild, before we get a clear picture of how they really compare.
My own sense is that Vera might have faster per-core performance, but nowhere near enough to win on either the basis of performance or efficiency, per socket. However, NVLink might mean that Vera scales up better. That could enable it to win on performance density. -
jeremyj_83 Reply
I don't think nVidia will equal Zen6 on per core performance. Vera was slightly faster per core than Zen5 depending on the benchmark. Odds are Zen6 will be at least that.bit_user said:My own sense is that Vera might have faster per-core performance, but nowhere near enough to win on either the basis of performance or efficiency, per socket. However, NVLink might mean that Vera scales up better. That could enable it to win on performance density. -
bit_user Reply
We'll see. I think it's not implausible. Especially when you take into account how much more bandwidth it has per core. On bandwidth-intensive stuff, Vera should easily pull ahead. But, that's also probably the minority of workloads. So, where the average falls, I'm not really sure.jeremyj_83 said:I don't think nVidia will equal Zen6 on per core performance.