Mastodon Skip to content
pulseofnations. Real News. Global Impact.
Subscribe
live markets
S&P 5007,677.28▲ 3.58%NASDAQ26,151.30▲ 4.71%DOW53,577.40▲ 3.14%GOLD4,684.30▲ 14.97%WTI80.08▼ 3.06%BRENT85.19▼ 3.59%EUR/USD1.1678▲ 2.65%USD/JPY158.94▼ 2.99%DXY98.95▼ 2.52%BTC$79,032▼ 1.19%ETH$2,465▼ 0.93%SOL$97.22▼ 3.23%TOTAL CRYPTO$2.68T▼ 3.39%

OpenAI Jalapeño Chip Beats Nvidia on Efficiency at Hot Chips 2026

OpenAI debuts benchmark results for its first custom inference chip, showing 1.5-1.9x more AI work per watt than Nvidia’s latest hardware.

Partner Surfshark VPN

OpenAI unveiled the first real-world performance results for Jalapeño, its custom-built inference chip, at the Hot Chips 2026 conference on Tuesday, showing the silicon outpaces Nvidia’s current hardware on a per-watt basis across multiple open-source and proprietary models.

The chip, co-developed with Broadcom and fabricated on TSMC’s N3P process, was designed from the ground up for one job: running large language model inference as fast and as efficiently as possible. On SemiAnalysis’ public InferenceX benchmark, Jalapeño sits on the Pareto frontier at 700 watts against Nvidia’s GB200 at 1.2 kilowatts and the GB300 at 1.4 kilowatts, meaning it delivers more useful throughput per unit of power at equal or better latency.

Benchmark Numbers Across Multiple Models

OpenAI tested Jalapeño on three workloads that are not co-designed for the chip: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Across all three, the chip delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. On Kimi K2.5, the largest model tested, Jalapeño achieved approximately 1.5 times higher peak mixed-token rate per kilowatt and 3.4 times lower latency than the GB300.

The hardware specifications back up those numbers. Jalapeño delivers 13.4 PFLOPs of MXFP4 matrix compute, pairs 15.4 terabytes per second of HBM4 bandwidth across 216 GiB of memory, and scales to 27 EFLOPs and 432 TiB across a 2,048-chip system. The network is split at 600 GB/s for a local 128-ASIC domain and 200 GB/s for the global domain, keeping the entire workload within one connected system to minimize data movement.

Crucially, Jalapeño uses single-token prediction while the Nvidia baselines use multi-token prediction with speculative decoding. Even in that direct comparison against the GB300 in multi-token mode on DeepSeek R1, single-token Jalapeño still leads with roughly 1.5 times higher peak mixed-token rate per kilowatt and 2.2 times lower end-to-end latency.

From Concept to Silicon in 16 Months

The project timeline was remarkably fast by semiconductor standards. An architecture concept appeared in late 2024, RTL froze in 2025, and the chip taped out by late 2025. Codex was running on it in early 2026, and ChatGPT followed not long after. Richard Ho, OpenAI’s head of hardware, told reporters that the results represent a very significant performance advance over the state of the art.

OpenAI also disclosed that it used its own AI models to help design the chip and to optimize the software that runs on it. Using Codex with GPT-Astra, the team brought three open-weight models that were not part of the original production plan to high performance within two months of receiving the silicon. For selected attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-expert-written implementations.

“The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly.”

Implications for the AI Chip Race

The announcement marks OpenAI’s formal entry into the custom silicon competition, a space previously dominated by Nvidia with AMD and Intel as distant challengers. OpenAI framed Jalapeño as evidence of a broader full-stack advantage, saying it can design models, products, serving software, chips, memory, networking, and systems together, using what it learns from real workloads to improve every layer.

OpenAI plans to begin deploying Jalapeño within its compute infrastructure by the end of 2026. The company said Gen 2 is deep in development and Gen 3 is taking shape, positioning this as the beginning of a multigenerational roadmap rather than a one-off experiment. The push comes as Nvidia reports fiscal Q2 earnings Wednesday after market close, with Wall Street expecting $92 billion in revenue and $2.09 EPS as Blackwell chip demand surges.

The competitive pressure is already reshaping the industry. Meta unveiled its MTIA 300 training chip last week, and Alphabet has long built custom TPUs for its own workloads. But OpenAI’s move is different: it is a model company turning itself into a hardware company, betting that controlling the silicon will let it serve AI workloads cheaper and faster than any third-party chip can. Whether that bet pays off at scale will become clearer once Jalapeño enters production deployment in the months ahead.

SourcesOpenAI; SemiAnalysis; TechCrunch; ServeTheHome; Hot Chips 2026
React to this dispatch
Share this dispatch X WhatsApp Bluesky Report an error
Written by

Founder and editor of Pulse of Nations, an independent wire service covering war, geopolitics, markets and technology.

discussion

Leave a Reply

Next dispatch OpenAI Cuts GPT-5.6 Sol Pricing Over 20% for Three Months Read →