THE CLAIMS
● 1.5 to 1.9x more work per watt than NVIDIA GB200 and GB300, and 1.7 to 3.6x lower latency.
● 2.1 to 4.1x on interactive workloads.
● Who ran the tests: OpenAI. SemiAnalysis verified the runs in person but did not execute the full suite.
● Timing: published one day before NVIDIA reports quarterly earnings.
What was announced
OpenAI presented the first published benchmark results for Jalapeño, its custom inference chip built with Broadcom, at the Hot Chips conference on 25 August. Testing ran on the InferenceX platform from SemiAnalysis across three models: GPT-OSS 120B, DeepSeek R1, and the 1-trillion-parameter version of Kimi K2.5.
| Measure |
Claimed advantage over GB200 / GB300 |
| Work per watt |
1.5x to 1.9x |
| Latency |
1.7x to 3.6x lower |
| Interactive workloads |
2.1x to 4.1x |
| Package TDP |
700W against 1,400W. Measured sustained power stayed at or below 550W |
Richard Ho, OpenAI's hardware chief, described it as a very significant performance advance over state of the art, with the payoff being that Jalapeño can serve more work per unit of power while returning responses more quickly. OpenAI credits a full-stack design that minimises data movement and keeps model state local.
Four things the headline numbers omit
1. OPENAI RAN THE TESTS
All figures come from OpenAI's own testing. SemiAnalysis verified the runs in person but did not execute the full suite itself. That is better than a pure vendor claim and short of independent benchmarking.
2. It cannot train models. Jalapeño is inference-only. Training is the workload where NVIDIA remains genuinely unchallenged, and nothing here touches it.
3. It was not tested against Vera Rubin. That is NVIDIA's next-generation platform — and the one slated to power the first gigawatt of NVIDIA systems that OpenAI itself agreed to deploy in the second half of 2026. Jalapeño beat the previous generation.
4. THE COMPARISON RAN SINGLE-TOKEN PREDICTION
The major comparisons pitted Jalapeño's single-token prediction against GB300 configurations doing the same — even though NVIDIA deployments commonly use multi-token prediction in production.
Against a GB300 running multi-token prediction, the peak efficiency lead shrinks to roughly 1.5x. An appendix comparison using all-in utility power — 1.18kW for Jalapeño against 2.55kW for the GB300 — also produces narrower gaps than the headline figures.
None of that makes the result fake. It makes it narrower than 4.1x suggests, and the honest range is closer to 1.5x on a like-for-like production configuration.
The timing
The benchmarks landed one day before NVIDIA reports quarterly earnings, with analysts expecting revenue near $92.27 billion. Draw your own conclusion about the scheduling. Worth stating plainly rather than pretending it is a coincidence, and equally worth not over-reading — Hot Chips sets its own calendar.
What is genuinely notable
Nine months from design to tape-out. OpenAI used its own models to accelerate development, which is the detail with the longest implications. A company shipping custom silicon that fast, partly designed by its own AI, changes what the barrier to entry looks like for everyone else.
Ho says a second-generation chip is deep into development with final design stage expected within months, and a third is already underway.
What it means for you
| If you are... |
The useful read |
| Building on the OpenAI API |
Nothing changes now. Deployment is small volumes by end of 2026, ramping through 2027 |
| Hoping to rent one |
You cannot. No rental market, no instance type, no plan to sell it externally |
| Watching inference pricing |
This is the mechanism behind cuts like GPT-5.6 Sol to $4/$20. Cheaper serving eventually reaches rate cards |
| Reading this as bad news for NVIDIA |
OpenAI says it will keep relying on NVIDIA GPUs alongside its own silicon, and still cannot train on Jalapeño |
FAQ
How much faster is Jalapeño than NVIDIA?
OpenAI claims 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower latency than GB200 and GB300, rising to 2.1 to 4.1 times on interactive workloads. Against a GB300 running multi-token prediction, the efficiency lead falls to roughly 1.5 times.
Who verified the benchmarks?
OpenAI ran them on SemiAnalysis's InferenceX platform. SemiAnalysis verified the runs in person but did not execute the full suite independently.
Can Jalapeño train models?
No. It is inference-only. Training remains the workload where NVIDIA hardware is unchallenged.
Was it compared against NVIDIA's latest chips?
Not against Vera Rubin, NVIDIA's next-generation platform — which is the one powering the first gigawatt of NVIDIA systems OpenAI agreed to deploy in the second half of 2026.
Can I use or rent a Jalapeño?
No. There is no rental market, no instance type, and no plan to sell the chip to other companies.
When does it ship?
Small volumes by the end of 2026, with a more significant ramp through 2027. OpenAI has not said how many it plans to deploy.
Does this mean OpenAI is dropping NVIDIA?
No. OpenAI says it will continue relying on NVIDIA GPUs alongside its own silicon, and it has separately agreed to deploy a gigawatt of NVIDIA systems this half.