by

OpenAI’s Jalapeño benchmarks beat Nvidia GB300

OpenAI has published its own Jalapeño benchmark numbers, claiming its 700W inference ASIC can outpace Nvidia’s flagship GB300 systems.

According to Tom’s Hardware, OpenAI arrived at Hot Chips with benchmark results showing that Jalapeño, its first in-house inference chip, delivers 1.5x to 1.9x higher throughput per kilowatt than Nvidia’s GB200 and GB300 rack systems.

The chip was co-developed with Broadcom and is aimed at inference, not training. Nvidia still owns the training market, and OpenAI is not claiming Jalapeño can wander into that fight and nick the crown.

OpenAI’s figures say Jalapeño delivers 1.7x to 3.6x lower end-to-end latency on SemiAnalysis’s public InferenceX suite. The tests used GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI’s one-trillion-parameter Kimi K2.5 model.

These are OpenAI’s own benchmark claims, run with SemiAnalysis engineers in OpenAI’s lab, rather than an independent industry shootout where everyone turns up with the same spanners and no marketing department watching.

OpenAI said Jalapeño did best at low-latency operating points. It claimed 8.6x to 104.3x more throughput per kilowatt at the GB300’s fastest previous time-between-tokens settings, which is the sort of range that should make readers hunt for the footnotes before fainting.

The comparison used published package TDPs, with Jalapeño at 700W against Nvidia accelerators rated at 1,200W and 1,400W. OpenAI said Jalapeño’s measured sustained power stayed at or below 550W during testing.

An appendix that uses all-in utility power presents a less heroic picture. That comparison puts Jalapeño at 1.18kW against 2.55kW for GB300, while GB300 with multi-token prediction narrows OpenAI’s peak efficiency lead to roughly 1.5x.

Nvidia’s Vera Rubin, which is meant to power the first gigawatt of Nvidia systems OpenAI agreed to deploy in the second half of 2026, was not included.

The main comparisons used Jalapeño single-token prediction against GB300 doing the same. Tom’s Hardware notes that Nvidia deployments commonly use multi-token prediction in production, which means the headline win is not quite the clean bar-room knockout it first appears to be.

SemiAnalysis described Jalapeño as “beating every Nvidia, AMD, and Google chip we have been able to test.” That is punchy, although it still sits inside a test regime arranged around OpenAI’s first public showing.

Each Jalapeño package pairs a compute die with six HBM4 stacks, giving it 216 GiB at 15.4 TB/s. Nvidia’s GB300 carries 288GB of HBM3E at a 1,400W rating, so OpenAI is making a memory-efficiency argument as much as a raw-speed one.

Scaling Jalapeño across OpenAI’s 10GW deployment agreement with Broadcom would make it a serious new claimant on HBM4 supply. That means OpenAI is not escaping the same wafer, memory and advanced packaging bottlenecks that make everyone else look twitchy.

A second-generation Jalapeño chip is close to tapeout, with work on a third generation already underway. The first part is said to use a TSMC 3nm-class process, keeping OpenAI in the same crowded manufacturing lanes as Blackwell and Rubin. The awkward bit is that OpenAI is benchmarking against a company it still badly needs. On 17 August, Nvidia agreed to provide up to $105 billion in financing for an OpenAI-leased data centre campus in Ohio.

OpenAI vice president of hardware Richard Ho said: “Nvidia is a really good partner, and we continue to need a lot of Nvidia.”

 

TOPICS:
AI inference  ·  broadcom  ·  custom ASIC  ·  gb300  ·  hbm4  ·  hot chips  ·  Jalapeño  ·  Nvidia  ·  openai

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews