256 Zen 6c cores make throughput look very different

AMD has expanded its performance data for the EPYC 9006 Venice family. In SPEC CPU 2026 intrate, a throughput-oriented workload that runs many copies of an application simultaneously, AMD says the 256-core EPYC 9996 delivers 2.24 times the performance of Nvidia’s 88-core Vera CPU.

The same data puts the 9996 at 2.37 times the score of Intel’s 128-core Xeon 6980P. Against AMD’s own 192-core EPYC 9965, the new flagship shows an approximately 78% generational uplift.

That test is naturally friendly to an architecture with enormous core density. SPEC Rate measures aggregate throughput, and Venice Dense brings 256 Zen 6c cores to a socket.

AMD uses 96 cores when it wants to talk about each core

For its per-core comparison, AMD changes the Venice configuration. A high-frequency 96-core setup is compared against Vera’s 88 Olympus cores.

AMD claims roughly a 20% per-core advantage in SPEC CPU 2026. That is a more interesting result than the raw 256-versus-88-core throughput chart because core count alone cannot explain it.

Power configuration deserves attention, though. Some of AMD’s comparison testing gives the down-cored Venice configuration access to a 600W envelope, while the commercial high-frequency 96-core part is specified below that level. It does not invalidate the measurement, but it changes what the graph represents.

Different GCC versions complicate the comparison

The biggest qualification is software. Nvidia published Vera results using GCC 15.2. Some of AMD’s newer Venice comparisons use GCC 16.1.

There is a sensible reason for that choice: GCC 16 contains newer optimization work and Zen 6 support. It also means the CPUs are not being compared with identical compiler releases.

Compiler changes can alter benchmark performance because they change the generated binaries. The size and even direction of that effect depend on the workload and compiler flags, which is precisely why a strict cross-architecture benchmark normally tries to keep the toolchain constant.

Venice also challenges Vera on memory bandwidth

AMD includes STREAM Triad results showing 1,298 GB/s of aggregate memory bandwidth from its Venice configuration versus 1,099 GB/s for Vera, an advantage of about 18%.

AMD calculates an approximately 8% per-core lead. This is another comparison between very different memory designs: Venice uses a 16-channel DDR5 platform, while Vera combines its 88-core CPU with high-bandwidth LPDDR5X through replaceable SOCAMM modules.

Memory is central to Nvidia’s pitch for Vera. The CPU is designed for agentic AI workloads where code execution, tool use, sandboxing, data processing and orchestration can keep the host processor busy even when the primary model computation lives on GPUs.

Venice is a family, not just a 256-core monster

The EPYC 9996 uses dense Zen 6c cores to reach 256 cores and 512 threads. AMD also has conventional Venice configurations scaling to 128 cores, high-frequency variants topping out at 96 cores and Venice-X models using 3D V-Cache.

The flagship offers up to 1,024 MB of L3 cache, while the new server platform supports 16 memory channels. Venice is manufactured on TSMC’s N2 process and uses a substantially different packaging approach from earlier EPYC generations.

The segmentation is deliberate. AMD can chase maximum socket throughput with Zen 6c while keeping fewer, faster Zen 6 cores for workloads where individual thread performance matters more.

Agentic AI still contains plenty of ordinary server work

AMD groups several of its new tests under “agentic AI workload performance.” Underneath the label are familiar components including NGINX, FAISS, SQL databases, TPCx-AI and multi-agent execution.

The terminology is new; much of the infrastructure work is not. An AI agent that runs code, queries databases, retrieves vectors and coordinates multiple tools can generate substantial CPU demand around the GPU inference itself.

That is why Vera and Venice are increasingly being framed as AI processors even though neither is replacing the accelerator. The host still has to run databases, runtimes, sandboxes, web services and orchestration.

The generational comparison may be the cleaner story

Cross-vendor results inevitably involve different architectures, memory systems and software stacks. AMD’s comparison against its own previous generation removes at least some of those variables.

The company reports roughly 78% more SPEC CPU 2026 intrate throughput from the 256-core EPYC 9996 than from the 192-core EPYC 9965. Its cloud, SQL, Java, cryptography and HPC slides also show substantial gains over Turin.

Independent testing will be needed to establish how those results translate into production servers with normalized power limits, compiler versions and software configurations. AMD has shown where it expects Venice to land. The cleaner Vera-versus-Venice answer still requires both CPUs on the same bench.