JC Davis

Exploded in 3D

CPU · GPU · DPU, taken apart

Drag to orbit. Hover any block. Click to pin it. Compare component classes below.

Explain
View
CPU
GPU
DPU
Step 1 of 4
CPU

GPU

DPU

62%
Compare:

Tip: pick a component class to spotlight it on all three chips at once — the cards below explain how each chip builds it, and why they differ.

WebGL isn't available in this browser, so the 3D stage can't render.
The comparison cards and spec tables below still work.

Scale

From one box to a data center

The same three chips, wired four ways. Watch the links change as you zoom out — and notice the moment the network stops being plumbing and becomes the computer.

WebGL isn't available, so the scale diagram can't render.
The captions still explain each topology.

Component diff

Compare

What it is

CPU

AMD EPYC 9654 “Genoa”

GPU

NVIDIA H100 SXM5 “GH100”

DPU

NVIDIA BlueField-3

At a glance

The same seven component classes, one row each. The pattern: every chip spends its transistor budget on the thing its job demands most.

Component classCPU — EPYC 9654GPU — H100DPU — BlueField-3
⚙ Compute cores 96× Zen 4general-purpose, branchy code 132 SMs · 16,896 CUDAmassively parallel math 16× Arm A78 + 16 DPAcontrol plane + packet I/O
🗄 Cache hierarchy 96 MB L2 + 384 MB L3hide DRAM latency for random access 50 MB shared L2rendezvous for 132 SMs, streaming data modest L2/L3 on meshpackets stream through, not cached
🧠 Memory system 12-ch DDR5, 461 GB/shuge capacity (TBs), moderate speed 5× HBM3, 3.35 TB/sbandwidth at all costs; 80 GB is enough 1 DDR5 ctrl, 16–32 GBjust a control plane; packets bypass RAM
🔗 Interconnect Infinity Fabric via sIODglue 12 chiplets into one CPU CoWoS interposer + NVLink 900 GB/sfeed SMs; scale to 256 GPUs PCIe Gen5 switch (32 lanes)own devices on the bus, bypass host
🌐 Network I/O — none on diethat's the DPU's job — none on dieNVLink is GPU-to-GPU, not network ConnectX-7 · 400 GbEthe whole point of the chip
🚀 Accelerators AVX-512 in-corevector math, no extra blocks 528 4th-gen Tensor Coresthe chip IS the accelerator crypto · RegEx · NVMe-oFinfrastructure tax, in fixed logic
📦 Package LGA-6096, socketedfield-replaceable server CPU SXM5 mezzanine, 700 Wpower + NVLink no slot can carry BGA card, air-cooledcheap, permanent, one per server

Provenance

Models are schematic — block sizes show hierarchy, not exact floorplans. Every headline number was verified against vendor or press sources, Sep 2026; anything estimated is labeled as an estimate.

  • AMD EPYC 9654 “Genoa” — 96C/192T Zen 4 (12 CCDs × 8 cores), 1 MB L2/core, 32 MB L3/CCD (384 MB), 5 nm CCDs + 6 nm 397 mm² sIOD, 12-ch DDR5-4800 (460.8 GB/s), 128× PCIe Gen5, SP5 LGA-6096, 360 W. TechPowerUp specs; CCD/CCX layout via Tom's Hardware.
  • NVIDIA H100 SXM5 “GH100” — full die 144 SMs (132 enabled), 128 CUDA + 4 Tensor cores/SM, 50 MB L2, 5 active HBM3 stacks of 6 on package, 5120-bit, 80 GB, ~3.35 TB/s, 80B transistors, 814 mm² TSMC 4N, 4th-gen NVLink 900 GB/s, 700 W. TechPowerUp specs; TweakTown; HWCooling (5-of-6 stacks).
  • NVIDIA BlueField-3 — 22B transistors, 16× Arm A78 v8.2 on coherent mesh, ConnectX-7 400 GbE/NDR, 4× crypto acceleration, 16 DPA packet cores, PCIe Gen5 switch, 1 DDR5 controller (16/32 GB on-card), inline IPsec/TLS, RegEx, NVMe-oF offload. NVIDIA technical blog; NVIDIA docs.
  • Exploded-stack geometry, pad counts, and internal block placement are estimates for illustration — vendors don't publish package CAD. Variant geometry (tile counts, die placement, V-Cache tiers) follows vendor block diagrams schematically, not to scale.
  • EPYC 9684X “Genoa-X” — 96C/192T, 12 CCDs, 96 MB L3/CCD (1,152 MB total) via 3D V-Cache stacking, same SP5 platform. AMD EPYC 9004X series product brief.
  • EPYC 9754 “Bergamo” — 128C/256T, 8× 16-core Zen 4c CCDs, 256 MB L3, 12-ch DDR5-4800 (460.8 GB/s), SP5. AMD EPYC 9004 series product brief.
  • Xeon 6980P “Granite Rapids” — 128C/256T, 504 MB L3, 3 compute tiles + 2 SoC tiles on EMIB, 12-ch DDR5-6400, 96× PCIe Gen5/CXL, LGA-7529, 500 W. Intel Xeon 6 product brief.
  • Apple M4 Max — up to 16 CPU cores, 40 GPU cores, 16-core Neural Engine, 128 GB unified LPDDR5X, 546 GB/s. Apple M4 Max tech specs.
  • Ryzen 9 9950X — 16C/32T Zen 5 (Granite Ridge), 4.3/5.7 GHz, 64 MB L3 + 16 MB L2, 2-ch DDR5, PCIe Gen5, 2-CU RDNA2 iGPU, AM5 LGA-1718, 170 W. TechPowerUp specs.
  • H100 NVL — paired PCIe cards joined by 3 NVLink bridges, 94 GB HBM3 per card (188 GB across the pair), ~3.9 TB/s per card, 350–400 W per card. NVIDIA H100 NVL product brief.
  • B200 “Blackwell” — 2 dies, 8× HBM3E stacks, 192 GB, ~8 TB/s, 10 TB/s NV-HBI die link, NVLink 5 at 1.8 TB/s. NVIDIA Blackwell architecture brief.
  • RTX 4090 — AD102, 128 SMs / 16,384 CUDA, 24 GB GDDR6X on 384-bit (~1.01 TB/s), PCIe 4.0 ×16, 450 W. NVIDIA GeForce RTX 4090 specs; TechPowerUp GPU database.
  • Instinct MI300X — 8× 5 nm XCD chiplets + 4× 6 nm I/O dies, 8 HBM3 stacks, 192 GB, 5.3 TB/s, 256 MB Infinity Cache, 750 W OAM. Tom's Hardware; AMD Advancing AI keynote (PDF).
  • BlueField-2 — 8× Arm A72, ConnectX-6 Dx (up to 200 Gb/s), PCIe Gen4, fixed-function accelerators. NVIDIA BlueField-2 product brief.
  • Intel IPU E2100 “Mount Evans” — up to 16× Arm Neoverse N1, 32 MB system cache, 48 GB LPDDR4x, PCIe 4.0 ×16, 2× 200 GbE, P4-programmable pipeline. Intel IPU E2100 product brief.
  • Pensando Elba — 7 nm, 16× Arm A72, 144 P4 MPUs, 2× 200 GbE, PCIe Gen4, DDR5-5600. AMD Pensando Elba product brief.
  • Marvell OCTEON 10 — 5 nm, up to 24× Arm Neoverse N2 (Armv9 + SVE2), DDR5, PCIe Gen5, up to 400 GbE, integrated 1 Tb switch, inline AI. Marvell OCTEON 10 product brief.

CPU vs GPU vs DPU, exploded. Schematic models; specs verified Sep 2026.