OpenAI's Jalapeño chip beats Nvidia Blackwell in early lab tests

OpenAI's Jalapeño chip beats Nvidia Blackwell in early lab tests

OpenAI has been developing its own inference chip, code-named Jalapeño, with partner Broadcom since the middle of 2024, going from initial team hiring to manufacturing tape-out in about 16 months, a fast timeline for an ASIC program. The chip program was first unveiled publicly in June, and Jalapeño itself was detailed at the Hot Chips conference. Analyst firm SemiAnalysis says OpenAI invited it into the company's labs to inspect the chip and run its InferenceX benchmark suite alongside OpenAI engineers, and reports that Jalapeño beats every Nvidia, AMD and Google chip it has tested on multiple open source models, measured in tokens generated per watt of power.

SemiAnalysis stresses that the headline framing, that Jalapeño is a chip specialized for OpenAI's own models, is wrong: it says the chip is a generalized inference accelerator designed to run all kinds of models and workloads well, rather than being tuned to one narrow case. In one of the tests, the chip hit over 700 tokens per second per user at concurrency 1 on the DeepSeek R1 model. On GPT-OSS it ran at approximately 1,400 tokens per second per user, and on Kimi K2.5 (the base model behind Cursor Composer 2.5) it reached nearly 700 tokens per second per user, more than 9 times the next best chip's throughput at 100 tokens per second per user. On GPT-OSS, Jalapeño's throughput per megawatt at matched interactivity was nearly double Nvidia's GB200 at its best point and more than 50 times GB200's result at concurrency 1. All of this was achieved using single-token prediction, without speculative decoding or splitting prefill and decode work across separate chip pools, an architectural choice OpenAI made because the mix of workloads changes over time.

SemiAnalysis flags several caveats. Every number in the comparison came from OpenAI; SemiAnalysis verified the InferenceX runs in person but did not run the full InferenceX suite itself and has not seen results on AgentX, its own benchmark for long-context, multi-turn workloads that it considers closer to real production traffic. It also argues the comparison to Nvidia's Blackwell is somewhat unfair, since Jalapeño is really competing against Nvidia's newer Rubin generation, which also uses HBM4 memory and is already shipping to customers, while Jalapeño remains at the engineering-sample stage. On a performance-per-total-cost-of-ownership basis, SemiAnalysis says Jalapeño and Vera Rubin are roughly head to head, even though Vera Rubin's figures benefit from speculative decoding (which can cut cost per token by roughly 3 to 5 times) while Jalapeño's do not yet use it.

On the hardware side, the tested A0 stepping of Jalapeño is only 9 months into the program; a further-optimized B0 stepping is already in fabrication and is expected to deliver about a 25% perf-per-watt improvement over A0. B0 offers 13.4 PFLOPs of MXFP4 compute on a single reticle-sized die built on TSMC's N3P process, against 17.5 PFLOPs of dense NVFP4 on a comparably sized Rubin die on the same node, while drawing only 700W versus Rubin's 900 to 1,150W per compute die. Jalapeño ships with HBM4 memory, delivering 15.4 terabytes per second of bandwidth per package, ahead of accelerators still using HBM3E. Off-package I/O runs over an N3E chiplet with 32 lanes of 800G SerDes, split between 24 lanes for scale-up within a rack and 8 lanes for scale-up across a 2,048-chip multi-rack domain, with PCIe Gen 5 connecting to the host CPU.

Key facts

  • OpenAI built Jalapeño with Broadcom, going from initial hiring to tape-out in about 16 months, and detailed it at Hot Chips after SemiAnalysis benchmarked it in OpenAI's labs.
  • On tokens generated per watt, Jalapeño beat every Nvidia, AMD and Google chip SemiAnalysis tested, reaching over 700 tokens per second per user on DeepSeek R1 and about 1,400 on GPT-OSS.
  • All figures came from OpenAI; SemiAnalysis verified the InferenceX runs in person but did not run the full suite and has not seen results on AgentX, its long-context benchmark.
  • SemiAnalysis says the fairer comparison is against Nvidia's newer, already-shipping Rubin chips rather than Blackwell, and that Jalapeño and Vera Rubin come out roughly even on cost per token.
  • The tested A0 chip is 9 months into the program and exists only as engineering samples; a B0 stepping already in fabrication is expected to add about 25% more performance per watt.

Why it matters

Jalapeño is OpenAI's first public proof that it can design competitive AI silicon on its own, rather than depending entirely on Nvidia. Going from hiring to tape-out in about 16 months is fast for an ASIC program, and the resulting chip reportedly leads Nvidia, AMD and Google hardware on tokens per watt across the models tested. Since OpenAI says it is currently limited by data center power rather than budget or floor space, tokens per watt, not raw FLOPs, is the metric that determines how much useful inference OpenAI can run per megawatt of power it can secure.

Who it affects

The result concerns OpenAI, its chip partner Broadcom, and the incumbents whose margins depend on OpenAI staying a customer: Nvidia, AMD and Google's TPU program. It also matters to Nvidia's newer Rubin line, which SemiAnalysis treats as Jalapeño's real competitor, and to other large AI labs and cloud operators watching whether in-house AI silicon can actually work, given that comparable programs at Meta and Microsoft have struggled to get off the ground.

How to use it

Jalapeño is not a product anyone can buy. The benchmarked A0 stepping is 9 months into the program and described as an engineering sample; a B0 stepping with roughly 25% better performance per watt is already in fabrication. No shipping date, general availability timeline or pricing is given, so for now Jalapeño is internal OpenAI infrastructure, not something outside developers or customers can access.

How solid is it

The numbers rest on OpenAI's own data. SemiAnalysis verified the InferenceX benchmark runs in person in OpenAI's labs but did not run the full InferenceX suite itself, and has not seen results on AgentX, its benchmark for long-context, multi-turn workloads that it considers more representative of production traffic than the 8k1k tests used here. SemiAnalysis itself calls the Blackwell comparison somewhat incomplete and unfair, arguing Jalapeño should be measured against Nvidia's newer Rubin chips instead, against which it says Jalapeño and Vera Rubin land roughly even on cost per token once Vera Rubin's use of speculative decoding is accounted for.

Risks and caveats

The comparison favors Jalapeño in ways the source itself flags: the tests use easier 8k1k workloads rather than AgentX's long-context, multi-turn scenarios that stress routers and cache management; Jalapeño's results use single-token prediction without speculative decoding, while the cited Vera Rubin figures use speculative decoding, which alone can cut cost per token by roughly 3 to 5 times; and the models tested (DeepSeek R1, Kimi K2.5, GPT-OSS) are smaller than the newest frontier models like DeepSeek V4 Pro and Kimi K3 that Nvidia and AMD have already benchmarked on AgentX. Jalapeño remains an engineering sample with no disclosed shipping timeline or pricing.

“If you have 1 gigawatt of power, then throughput per watt is revenue”

— Jensen, speaking at Computex 2026