CHIPS·WARS 2026 — LOADING SILICON…
3D SCENES · INTERACTIVE CHARTS · BILINGUAL EN/ES
◈ Special report · August 25, 2026 · Updated with primary sources

JalapeñovsRubinvsCerebras

On August 25, 2026, OpenAI published the first benchmarks of Jalapeño, its first custom inference chip. The timing says everything: NVIDIA's Rubin has just entered full production, and Cerebras floated on Nasdaq on the promise of wafer-scale speed. Three radically different answers to one question — how should the world serve AI? — settled in silicon. Rotate the scene, click a chip, then scroll: we'll open each package layer by layer and show you why memory, not math, decides who wins inference.

700 W · 1.5–1.9× work/watt
3.6 EF FP4 · per NVL72 rack
43.2 PB/s · wafer SRAM bandwidth

🖱 Drag or two-finger rotate · click a chip
01 / THE THREE TITANS

Three philosophies of silicon

A reticle-sized 3nm ASIC designed by the world's largest AI consumer; the sixth generation of the ecosystem king; and an entire 300 mm wafer turned into a single processor. Three different bets on the same roofline — and three companies whose futures depend on them.

OpenAI · Broadcom · TSMC 3nm

Jalapeño 🌶️

Inference “Intelligence Processor” · Gen 1
700 W
TDP (≤550 W sustained)
9 mo
design → tape-out
UnveiledJun 24, 2026
Compute die [est.]~840 mm² · reticle-size
Memory6× HBM4 · 216 GiB · 15.4 TB/s
NetworkEthernet fabric · scale to 2,048 XPUs
Efficiency vs GB200/3001.5–1.9× work/watt
Interactive perf2.1–4.1× · latency 1.7–3.6× lower
RoadmapGen 2 deep in development · Gen 3 taking shape
Part of a 10 GW OpenAI–Broadcom accelerator program through 2029 — serving 800M+ weekly ChatGPT users.
NVIDIA · TSMC N3P + CoWoS-L

Rubin R200

Dual-die GPU · Vera Rubin NVL72 platform
50 PF
NVFP4 / package (sparse)
22 TB/s
HBM4 bandwidth (official)
Transistors336 B · 224 SMs · 896 tensor cores
Memory288 GB HBM4 (8× 12-Hi)
Package2 compute dies + 2 I/O tiles
Rack (NVL72/144)3.6 EF FP4 · 75 TB fast mem · NVLink 6
Vera CPU88 custom cores · 1.5 TB LPDDR5X
StatusFull production · clouds 2H 2026
$1 T in combined orders through 2027 (GTC 2026). Rubin Ultra follows in 2027.
Cerebras (Nasdaq: CBRS) · TSMC 5nm

WSE-3 Turbo

Wafer-Scale Engine · CS-4 rack system
4 T
transistors / wafer
4,400
tok/s/user (GPT-OSS 120B)
AI cores900,000 · ~0.05 mm² each
On-wafer SRAM44 GB @ 43.2 PB/s
Compute250 PFLOPS sparse FP16 /wafer
CS-4 rack3 wafers · 750 PF · 7.2 Tbps I/O
CompanyIPO May 2026 · ~$95–107 B
Inference claimup to 30× vs GPU systems
$10 B OpenAI capacity deal (Jan 2026) · $25.4 B in remaining obligations.
02 / INSIDE THE SILICON

3D teardowns — every piece, piece by piece

A chip is a stack of engineered layers. Drag to rotate, pull the slider to explode the stack, and click any part — in the scene or in the list — to learn what it is, what it costs, and why it decides performance. Scale is stylized; every number is sourced.

drag · explode slider · click parts
03 / 3D COMPARATOR

Scales that hurt

Log-scale 3D bars — these numbers differ by orders of magnitude, so linear bars would lie. Orbit the scene and switch metrics. Values marked “?” are undisclosed; estimates are flagged.

★ BONUS LAB / HEAD-TO-HEAD

Duel any two chips — live

Pick two contenders and the lab scores them head-to-head on every measured category, computing winners and ratios from the same sourced data as the rest of the page — then writes the verdict for you.

04 / THE MEMORY WALL

Why bandwidth — not FLOPS — decides inference

Generating a token at low batch size means reading every model weight once per token: a 70B model in FP16 streams ~140 GB to emit a single word, doing only ~2 FLOPs per parameter. That's ~1–2 FLOPs per byte moved, while modern accelerators need ~200–2,000 to saturate their math. The consequence: decode speed ≈ memory bandwidth. This section shows the theory, the hardware ladder, and lets you compute it yourself.

The roofline, live

Each chip's attainable performance = min(bandwidth × intensity, compute ceiling). Move the batch slider: at batch 1 every machine is bandwidth-bound — then watch Cerebras hit its ceiling first (its SRAM is so fast the math becomes the limit), while GPUs stay memory-starved until huge batches.
1

The memory ladder · log scale

Capacity vs bandwidth vs latency: you never win all three. SRAM is nanoseconds-fast but tiny; HBM4 balances; GDDR7/LPDDR5X trade bandwidth for capacity and cost. Log-scale bandwidth per chip/wafer.

Analogy (Berkeley CS 61C): SRAM is the book on your desk; HBM is the library stacks; LPDDR is interlibrary loan. Bigger shelves are always slower to reach.

KV-cache calculator — where does your model actually fit?

The KV cache grows with every token of context and can't be shared between users. Pick a model, drag the context, choose precision — and see which package can hold weights + cache. That's the real reason chip choice is a product decision.
KV cache
Weights
Total resident
users / package (mem-only)
🌶️ Jalapeño · 216 GiB HBM4
🏆 Rubin R200 · 288 GB HBM4
⚡ WSE-3 Turbo · 44 GB SRAM

Feel the speed — a token-rate simulator

Humans read at ~238 words/min (Brysbaert 2019) ≈ 5–6 tokens/s. Below that, AI feels slow; above ~50 tok/s the model types faster than you can read. Watch the same paragraph stream at each machine's real output rate.
tokens / second
this paragraph (s)
a 1,000-token answer (s)
how it feels
05 / TACTICAL RADAR

Strengths squared off

Qualitative 0–5 assessment from official data and claims (method: this page's sourced numbers; vendor marketing discounted). Tap a card to mute a contender.

06 / JBENCH · JALAPEÑO'S FIRST NUMBERS

First official results

Benchmarks published by OpenAI (SemiAnalysis InferenceX suite), Aug 25 2026: Jalapeño vs the NVIDIA GB200 (1,200 W) and GB300 (1,400 W) systems currently serving ChatGPT. Workload: 8k-token input / 1k-token output, mixed prefill+decode.

07 / DATAFLOW DIAGRAMS

How each machine thinks

The path of one inference request through each architecture. Animated lines are data in motion.

08 / TIMELINE

Seven years to this week

From the first trillion-transistor wafer (2019) to a three-way war for inference (2026) — and what's coming next.

09 / VOICES & ECONOMICS

What the principals are saying — and spending

10 / FULL SPEC SHEET

Master spec table

🌶️ OpenAI JalapeñoNVIDIA Rubin R200Cerebras WSE-3 Turbo
TypeInference-only ASIC (“Intelligence Processor”)Dual-die GPU + I/O tilesWafer-scale (wafer = chip)
UnveiledJun 24, 2026GTC Mar 2025 · production CES Jan 2026WSE-3 Turbo: Aug 2026
TransistorsUndisclosed (die ~840 mm² [est.])336 B4 trillion (4×10¹²)
ProcessTSMC 3nm-classTSMC N3P + N5 I/OTSMC N5
CoresUndisclosed224 SMs · 896 tensor cores900,000 dataflow cores
Peak computeNot published (metric: work/watt)50 PF NVFP4 sparse / 35 PF dense250 PF sparse FP16 /wafer
Memory6× HBM4 · 216 GiB · 15.4 TB/s288 GB HBM4 · up to 22 TB/s (+54 TB LPDDR5X rack)44 GB on-wafer SRAM · 43.2 PB/s
Power700 W TDP · ≤550 W sustainedUndisclosed (~1.8 kW est.)~40–54 kW/wafer [est.]
Rack / systemEthernet fabric → 2,048 XPUsNVL72: 3.6 EF FP4 · 75 TB · 260 TB/s NVLink 6CS-4: 3 wafers · 750 PF · 129.6 PB/s
Star metric1.5–1.9× work/watt vs GB200/3003.3× GB300 NVL72 (rack)4,400 tok/s/user · up to 30× vs GPU
EcosystemOpenAI stack · 3 open models in 2 monthsCUDA · every major cloudCerebras SDK · OpenAI API partner
AvailabilityDeploying late 2026 · small volumesClouds 2H 2026Shipping Q3 2026 (CS-4)
Next genGen 2 in development · Gen 3 taking shapeRubin Ultra 2027: 15 EF · 1 TB HBM4eCS-5/CS-6 racks already reusable
11 / VERDICT

Who wins what — and what nobody tells you

🌶️ Efficiency & interactive latency: Jalapeño

Wins work per watt: 1.5–1.9× over the Blackwell systems serving ChatGPT today, at 700 W vs 1,200–1,400 W. Up to 104.3× more throughput at the GPUs' previous latency (DeepSeek R1). Secret weapon: it only has to be perfect at one job — OpenAI's own models. The model is proven: Google's TPU v1 delivered 30–80× perf/W over 2015-era hardware (Jouppi et al., ISCA 2017).

🏆 Absolute scale & ecosystem: Rubin

Wins total capacity: 3.6 FP4 exaflops and 75 TB of fast memory per rack, NVLink 6 at 260 TB/s, CUDA behind it, ~$1 T of orders, and a 2027 successor (Rubin Ultra: 15 EF, 1 TB HBM4e) already taped to the roadmap. It trains and infers, and it's the only one of the three you can simply buy.

⚡ Pure per-user speed: Cerebras

Wins tokens per user: 4,400 tok/s (GPT-OSS 120B), 1,000+ tok/s on 10T-parameter models, 2 µs wafer-to-wafer hops — “in one second what a GPU rack needs 30 seconds for.” Limit: 44 GB of SRAM per wafer means weights stream from MemoryX, and trillion-parameter models span many wafers — the burden moves to the interconnect.

The uncomfortable truth

These benchmarks were run inside OpenAI's lab against last-generation GPUs — Vera Rubin, the actual 2026 competitor, was not in the comparison, and all-in wall power differs too (1.18 kW vs 2.55 kW per package — Tom's Hardware). OpenAI itself says it will “continue to widely deploy accelerators from NVIDIA and other partners,” and Richard Ho told Bloomberg: “Nvidia is a really good partner, and we continue to need a lot of Nvidia.” Meanwhile NVIDIA answers on both flanks — Rubin CPX for context-heavy prefill, and a Rubin Ultra refresh already scheduled for 2027. The likely 2027 picture isn't a winner — it's a portfolio: Rubin for scale, custom ASICs for cost, wafer-scale for latency.

12 / SOURCES

Sources & methods

Primary sources first (vendor posts, official specs, SEC-grade reporting), then specialist analysis. Pre-launch specs change; estimates are marked as such where used.