At a glance
- Benchmark per watt, not per chip. OpenAI tested Jalapeño on InferenceX across three models: Generative Pre-trained Transformer (GPT) Open Source Software (OSS) 120B, DeepSeek R1 670B, and Kimi K2.5 1T.
- Match parts to workloads. Jalapeño is built strictly to run models, not train them.
- Right-size before you re-buy. The cheapest upgrade is often the memory and silicon you already own.
- OpenAI’s custom chip, Jalapeño, did 1.5 to 1.9 times more work per watt than Nvidia’s GB300 on the August 25 InferenceX test.
OpenAI’s first custom chip beat Nvidia on the one metric that limits Artificial Intelligence (AI) scale: power use . The new Jalapeño chip runs models instead of training them. It does 1.5 to 1.9 times more work per watt than Nvidia’s GB300 . It also cuts wait times by up to 3.6 times. It draws under 550 watts, compared to 1,400 watts for the Nvidia part .
Buyers build their own chips
Why it matters: Power, not chip supply, now limits how fast companies can add server space . OpenAI’s hardware lead tracks success in tokens per megawatt, not tokens per second . A chip wins by doing more useful work per watt. It no longer wins just by hitting a high peak speed.
The big picture: Every major tech lab has made this same choice. OpenAI calls Jalapeño a strong path alongside hardware from Nvidia, Advanced Micro Devices (AMD), Broadcom, Cerebras, and CoreWeave . The largest buyers of merchant Graphics Processing Units (GPUs) are all building custom chips at once .
Google runs its Tensor Processing Unit (TPU) line. Amazon runs Trainium. Microsoft runs Maia. Meta’s Iris chip also entered production this month .
“The metric that matters is tokens generated per dollar per watt. Peak theoretical Tera Floating Point Operations Per Second (TFLOPS) on a spec sheet is vendor marketing.” — Founder Note
Jalapeño marks the first time a custom lab chip posted a public test against Nvidia’s top part . The design went from a first sketch to tape-out—the final step before making the chip—in just nine months . OpenAI even used its own Large Language Models (LLMs) to speed up parts of the design work .
The infrastructure playbook
You cannot buy a Jalapeño chip. It belongs to OpenAI and will not ship in high volume until 2027 or 2028 . But this test signals a three-part shift in how to judge hardware.
The honest catch
The catch: These are vendor-reported numbers from before the chips hit real servers . First production runs at the end of 2026 will be small. Large setups land in 2027, and a full scale-up extends into 2028 . OpenAI has also not promised to pass these power savings through to Application Programming Interface (API) pricing .
The chip is real, but Nvidia still sells the vast majority of hardware. The test proves custom silicon is closing the cost gap. However, it will not threaten Nvidia’s training lead for at least 18 months.
Go deeper: Explore the sources and internal links below.
-
Power limits scale, pushing massive buyers like OpenAI, Google, Amazon, Microsoft, and Meta to build their own custom silicon .
-
Server teams must track speed-per-watt and match specific parts to workloads, though Jalapeño will not ship at scale until 2027 .
References
[1] OpenAI, “OpenAI and Broadcom unveil LLM-optimized inference chip” (June 24, 2026).
[3] Sarah Friar, “The full stack behind abundant intelligence” (Aug 25, 2026).