The Information Machine
Following·since 25 Aug 2026·Day 5·8 sources·updated 28 Aug 2026

OpenAI's Jalapeño inference chip

The gist

OpenAI Releases Jalapeño Benchmarks Claiming Efficiency Lead Over Nvidia

OpenAI's first custom inference chip, if the benchmarks hold, substantially reduces its dependence on Nvidia for serving workloads. The chip also demonstrates AI-assisted hardware development, with design-to-tapeout completed in nine months and AI-generated kernels outperforming human-written code on selected blocks.

The full picture

OpenAI published benchmark results for Jalapeño, its first custom AI inference chip, on August 25, 2026, claiming 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency compared to leading commercially available AI systems. For highly interactive workloads, Jalapeño delivers 2.1 to 4.1 times higher performance. At a matched DeepSeek R1 decoding speed of 169.41 tokens per second per user, Jalapeño produced 12,258 mixed tokens per second per kilowatt versus 118 for Nvidia's GB300, a 104.3x ratio. The chip is rated at 700W, compared to 1,200W for the GB200 and 1,400W for the GB300. Built with Broadcom and fabricated by TSMC, Jalapeño was validated across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The chip is inference-only and cannot train models, so OpenAI continues to rely on Nvidia for training. OpenAI plans a small batch in 2026 with more in 2027, and two additional chip generations are in development.

How it developed
28 August 2026

Further coverage repeated benchmark claims of Jalapeño surpassing Nvidia's flagship on speed and efficiency

Coverage on August 27 of OpenAI's August 25 Jalapeño benchmarks specified performance gains for highly interactive workloads at 2.1 to 4.1 times, giving a lower bound absent from earlier reporting. The same coverage introduced a safety argument: an OpenAI executive identified as roon contended that inference at this speed creates threat models where misaligned frontier models could infiltrate systems faster than human responders can act, requiring autonomous detection and shutdown rather than monitoring alone.

27 August 2026

Analysis noted ultrafast inference creates new threat models requiring autonomous detection and shutdown

Coverage following the August 25 benchmark release confirmed Jalapeño was built in partnership with Broadcom and fabricated by TSMC. The chip handles inference only, leaving OpenAI still dependent on Nvidia for training workloads; a small batch of Jalapeño-powered systems is planned for 2026, more in 2027, and two additional chip generations are in development.

First citedMilk Road AI TwitterThe NeuronSemafor Technology
26 August 2026

Additional coverage confirmed inference-only limitation and noted Google, Amazon, Microsoft, and Anthropic are also developing custom chips

SemiAnalysis published an analysis August 26 extending OpenAI's Jalapeño chip benchmarks, released August 25, to Nvidia's upcoming Vera Rubin generation, finding Jalapeño beats Vera Rubin's published output-token throughput per megawatt. SemiAnalysis identified electricity as OpenAI's primary design constraint, described a networking architecture making thousands of chips function as a single inference machine, and suggested OpenAI may be designing for models with tens of trillions of parameters or million-token context windows.

25 August 2026

SemiAnalysis published analysis finding Jalapeño beats Nvidia Vera Rubin on output-token throughput per megawatt

19 August 2026

Supply chain checks show Nvidia holding over 80% of AI server market, with ASIC-based servers at 8-11% of shipments

Sources
3 more sources
Sam Altman Twitter (OpenAI)
Semafor Technology
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free