On August 25, 2026, OpenAI published the first benchmarks of Jalapeño, its first custom inference chip. The timing says everything: NVIDIA's Rubin has just entered full production, and Cerebras floated on Nasdaq on the promise of wafer-scale speed. Three radically different answers to one question — how should the world serve AI? — settled in silicon. Rotate the scene, click a chip, then scroll: we'll open each package layer by layer and show you why memory, not math, decides who wins inference.
A reticle-sized 3nm ASIC designed by the world's largest AI consumer; the sixth generation of the ecosystem king; and an entire 300 mm wafer turned into a single processor. Three different bets on the same roofline — and three companies whose futures depend on them.
A chip is a stack of engineered layers. Drag to rotate, pull the slider to explode the stack, and click any part — in the scene or in the list — to learn what it is, what it costs, and why it decides performance. Scale is stylized; every number is sourced.
Barras 3D en escala logarítmica — estos números difieren en órdenes de magnitud, así que las barras lineales mentirían. Rota la escena (arrastra; dos dedos en táctil) y cambia de métrica. Los valores “?” no están divulgados; los estimados están marcados.
Elige dos contendientes y el laboratorio los puntúa cara a cara en cada categoría medida, calculando ganadores y ratios con los mismos datos con fuente del resto de la página — y luego redacta el veredicto por ti.
Generating a token at low batch size means reading every model weight once per token: a 70B model in FP16 streams ~140 GB to emit a single word, doing only ~2 FLOPs per parameter. That's ~1–2 FLOPs per byte moved, while modern accelerators need ~200–2,000 to saturate their math. The consequence: decode speed ≈ memory bandwidth. This section shows the theory, the hardware ladder, and lets you compute it yourself.
Analogía (Berkeley CS 61C): la SRAM es el libro sobre tu escritorio; HBM son los estantes de la biblioteca; LPDDR es el préstamo interbibliotecario. Estantes más grandes siempre cuestan más alcanzar.
Valoración cualitativa 0–5 a partir de datos y afirmaciones oficiales (método: los números con fuente de esta página; el marketing de cada fabricante, descontado). Toca una tarjeta para silenciar a un contendiente.
Benchmarks published by OpenAI (SemiAnalysis InferenceX suite), Aug 25 2026: Jalapeño vs the NVIDIA GB200 (1,200 W) and GB300 (1,400 W) systems currently serving ChatGPT. Workload: 8k-token input / 1k-token output, mixed prefill+decode.
El camino de una petición de inferencia por cada arquitectura. Las líneas animadas son datos en movimiento.
De la primera oblea de un billón de transistores (2019) a una guerra a tres bandas por la inferencia (2026) — y lo que viene después.
| 🌶️ OpenAI Jalapeño | NVIDIA Rubin R200 | Cerebras WSE-3 Turbo | |
|---|---|---|---|
| Type | Inference-only ASIC (“Intelligence Processor”) | Dual-die GPU + I/O tiles | Wafer-scale (wafer = chip) |
| Unveiled | Jun 24, 2026 | GTC Mar 2025 · production CES Jan 2026 | WSE-3 Turbo: Aug 2026 |
| Transistors | Undisclosed (die ~840 mm² [est.]) | 336 B | 4 trillion (4×10¹²) |
| Process | TSMC 3nm-class | TSMC N3P + N5 I/O | TSMC N5 |
| Cores | Undisclosed | 224 SMs · 896 tensor cores | 900,000 dataflow cores |
| Peak compute | Not published (metric: work/watt) | 50 PF NVFP4 sparse / 35 PF dense | 250 PF sparse FP16 /wafer |
| Memory | 6× HBM4 · 216 GiB · 15.4 TB/s | 288 GB HBM4 · up to 22 TB/s (+54 TB LPDDR5X rack) | 44 GB on-wafer SRAM · 43.2 PB/s |
| Power | 700 W TDP · ≤550 W sustained | Undisclosed (~1.8 kW est.) | ~40–54 kW/wafer [est.] |
| Rack / system | Ethernet fabric → 2,048 XPUs | NVL72: 3.6 EF FP4 · 75 TB · 260 TB/s NVLink 6 | CS-4: 3 wafers · 750 PF · 129.6 PB/s |
| Star metric | 1.5–1.9× work/watt vs GB200/300 | 3.3× GB300 NVL72 (rack) | 4,400 tok/s/user · up to 30× vs GPU |
| Ecosystem | OpenAI stack · 3 open models in 2 months | CUDA · every major cloud | Cerebras SDK · OpenAI API partner |
| Availability | Deploying late 2026 · small volumes | Clouds 2H 2026 | Shipping Q3 2026 (CS-4) |
| Next gen | Gen 2 in development · Gen 3 taking shape | Rubin Ultra 2027: 15 EF · 1 TB HBM4e | CS-5/CS-6 racks already reusable |
Wins work per watt: 1.5–1.9× over the Blackwell systems serving ChatGPT today, at 700 W vs 1,200–1,400 W. Up to 104.3× more throughput at the GPUs' previous latency (DeepSeek R1). Secret weapon: it only has to be perfect at one job — OpenAI's own models. The model is proven: Google's TPU v1 delivered 30–80× perf/W over 2015-era hardware (Jouppi et al., ISCA 2017).
Wins total capacity: 3.6 FP4 exaflops and 75 TB of fast memory per rack, NVLink 6 at 260 TB/s, CUDA behind it, ~$1 T of orders, and a 2027 successor (Rubin Ultra: 15 EF, 1 TB HBM4e) already taped to the roadmap. It trains and infers, and it's the only one of the three you can simply buy.
Wins tokens per user: 4,400 tok/s (GPT-OSS 120B), 1,000+ tok/s on 10T-parameter models, 2 µs wafer-to-wafer hops — “in one second what a GPU rack needs 30 seconds for.” Limit: 44 GB of SRAM per wafer means weights stream from MemoryX, and trillion-parameter models span many wafers — the burden moves to the interconnect.
These benchmarks were run inside OpenAI's lab against last-generation GPUs — Vera Rubin, the actual 2026 competitor, was not in the comparison, and all-in wall power differs too (1.18 kW vs 2.55 kW per package — Tom's Hardware). OpenAI itself says it will “continue to widely deploy accelerators from NVIDIA and other partners,” and Richard Ho told Bloomberg: “Nvidia is a really good partner, and we continue to need a lot of Nvidia.” Meanwhile NVIDIA answers on both flanks — Rubin CPX for context-heavy prefill, and a Rubin Ultra refresh already scheduled for 2027. The likely 2027 picture isn't a winner — it's a portfolio: Rubin for scale, custom ASICs for cost, wafer-scale for latency.
Primero fuentes primarias (publicaciones de los fabricantes, specs oficiales, reportajes de nivel SEC), después análisis especializado. Las specs de pre-lanzamiento cambian; las cifras derivadas van marcadas [est.] donde se usan.