The Numbers
Against Nvidia's GB200 and GB300, OpenAI's Jalapeño inference ASIC delivers:
| Metric | Jalapeño vs Blackwell |
| Peak throughput per watt | 1.5x - 1.9x higher |
| End-to-end latency | 1.7x - 3.6x lower |
| Minimum time between tokens | 2.1x - 4.1x lower |
| Peak compute | Substantially lower |
Jalapeño reaches 3.4 MXFP8 PFLOPS and 13.4 MXFP4 PFLOPS. Blackwell's raw figures are higher, and OpenAI is not pretending otherwise.
Losing on Compute Is the Design
A chip that loses on peak FLOPS and wins on latency, time-between-tokens and performance per watt has not fallen short. It has been built for a different job.
Training is compute-bound. Inference is latency-bound. When someone is waiting on a response, what governs the experience is how fast the first token arrives and how steadily the rest follow - which is exactly what "2.1x to 4.1x lower minimum time between tokens" describes. Peak FLOPS is the number you optimise when you are training a model, and OpenAI is not training on this.
Performance per watt is the other half, and at inference scale it is the whole economic argument. A 1.5x to 1.9x advantage there compounds across every token served, forever, against a power bill that is currently the binding constraint on the entire industry.
Memory Is the Real Spec
- 216 GB of HBM4 per accelerator, at up to 15.4 TB/s bandwidth.
- A NUMA-style spatial architecture of 64 core slices, each with dedicated HBM.
- A 128-accelerator domain exceeds 1 PB/s aggregate bandwidth.
- At full 2,048-processor scale: 432 TB of HBM4 and 32 PB/s aggregate.
Pairing each core slice with its own memory is the structural choice behind the latency numbers. Inference is memory-bandwidth-starved far more often than it is compute-starved, and this layout attacks that directly.
It Hosts on AMD, Not Nvidia
Jalapeño deploys with AMD EPYC "Turin" CPUs as hosts.
Worth stating plainly: OpenAI built an accelerator to reduce its dependence on Nvidia, and then chose AMD for the host CPUs too. The result is a production inference stack with no Nvidia silicon in it.
Nine Months, and the Chip Helped Design Itself
| Milestone | Date |
| Design began | February 2025 |
| Tape-out | November 2025 |
| First silicon | May 2026 |
| Codex running on it | May 2026 |
Nine months from RTL to tape-out is fast for a reticle-sized accelerator, and the reason is the second disclosure: more than half the core was written in the XLS hardware language, with AI optimisation producing measurable gains on individual blocks - 56% on a BF16 multiplier, 21% on FP4 dot-product blocks, 10% on FP32 accumulators.
Those are block-level improvements rather than whole-chip ones, and the distinction matters. But a 56% improvement on a multiplier is a real engineering result, and the loop is now closed: models designed significant parts of the silicon that now serves those models.
Process node and die size were not disclosed.
What It Changes
- Inference economics. A 1.5-1.9x performance-per-watt advantage on the serving side is the kind of structural cost difference that shows up in API pricing eventually.
- Nvidia's position at inference specifically. Training demand is untouched. Serving is where a purpose-built ASIC can beat a general-purpose GPU, and now one demonstrably does.
- The design cycle. Nine months RTL-to-tape-out, with AI writing half the core, shortens the gap between deciding you need custom silicon and having it.
Background on the chip's unveiling with Broadcom is in our earlier piece.
FAQ
Is Jalapeño faster than Nvidia Blackwell?
On inference metrics, yes - 1.5x to 1.9x higher throughput per watt, 1.7x to 3.6x lower end-to-end latency, 2.1x to 4.1x lower minimum time between tokens. On raw peak compute, no. Blackwell is substantially higher.
What CPU does it use?
AMD EPYC "Turin" as the host CPU, not Nvidia. The production inference stack contains no Nvidia silicon.
How much memory does it have?
216 GB of HBM4 per accelerator at up to 15.4 TB/s. A full 2,048-processor configuration reaches 432 TB and 32 PB/s aggregate bandwidth.
Did AI design the chip?
Partly. More than half the core was written in the XLS hardware language, with AI optimisation delivering 56% improvement on a BF16 multiplier, 21% on FP4 dot-product blocks and 10% on FP32 accumulators. Humans ran the design.