AI

OpenAI’s Jalapeño chip runs inference at 1.9x Nvidia Blackwell’s efficiency per watt

Adrian Kessler

OpenAI built its own chip. The company presented Jalapeño at the Hot Chips conference on August 25, a silicon design developed in partnership with Broadcom that processes inference requests at 85,448 tokens per second per kilowatt. Nvidia‘s Blackwell architecture, the current market standard, reaches 44,960 on the same metric. The ratio is 1.9 times.

The metric matters because inference is how AI systems run in production. Training happens once; inference happens every time someone sends a query to ChatGPT, uses an API, or interacts with any application that runs a language model. The cost of inference at scale is the primary driver of AI service economics. A chip that cuts inference energy consumption roughly in half, if deployed widely, changes the per-query cost of running models.

On interactive workloads, where response latency is the limiting factor, Jalapeño’s advantage extends further. The SemiAnalysis report on the Hot Chips presentation put the range at 2.1 to 4.1 times better than Blackwell on those benchmarks. The spread reflects variation across specific task types; the lower bound is the more conservative comparison.

The chip handles inference only. It does not run the training workloads that produced the models Jalapeño is designed to run. Those still require Nvidia hardware. Broadcom manufactured the chip under a custom design agreement with OpenAI rather than selling it as a commercial product. The arrangement gives OpenAI a purpose-built inference layer without requiring it to compete in the merchant chip market.

OpenAI’s deployment timeline is limited. Small-scale rollout is planned before end of 2026; broader deployment is a 2027 goal. The chip remains in the engineering sample phase, which means the production version has not yet shipped at scale. The benchmarks at Hot Chips represent the designed target, not a verified production run. The gap between conference announcement and operational deployment for custom silicon is typically measured in years.

The announcement fits a broader shift among large AI companies toward custom silicon. Google has operated its own Tensor Processing Units in production for a decade. Amazon runs Trainium and Inferentia in AWS. Meta operates its own inference chips internally. OpenAI’s entry confirms that at the scale of hundreds of millions of active users, the economics of relying entirely on external GPU vendors become difficult enough to justify the cost of a custom design program.

Tags: , , , , ,

Discussion

There are 0 comments.