Mastodon Skip to content
pulseofnations. Real News. Global Impact.
Subscribe
live markets
S&P 5007,711.76▲ 3.81%NASDAQ26,402.42▲ 6.13%DOW53,559.99▲ 1.54%GOLD4,529.90▲ 12.23%WTI83.40▲ 5.22%BRENT88.10▲ 4.77%EUR/USD1.1587▲ 1.92%USD/JPY160.04▼ 2.28%DXY99.68▼ 1.68%BTC$78,128▲ 0.94%ETH$2,451▲ 0.99%SOL$105.01▲ 1.25%TOTAL CRYPTO$2.66T▼ 1.56%

OpenAI Jalapeno Chip Beats Nvidia on Efficiency Benchmarks

OpenAI’s first custom inference chip delivers 1.5 to 1.9 times more AI work per watt than Nvidia’s Blackwell in Hot Chips results

Partner Surfshark VPN

OpenAI has revealed the first benchmark results for Jalapeno, its custom AI inference chip developed with Broadcom, and the numbers show it outperforming Nvidia Blackwell systems across every key efficiency metric. The chip, fabricated on TSMC 3-nanometer process, was tested on SemiAnalysis public InferenceX benchmark and posted results that surprised industry observers who expected a first-generation custom silicon effort to lag behind established players. Across three public models tested, Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, the performance advantage widened to 2.1 to 4.1 times higher than existing hardware. “The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI head of hardware, during a press call following the presentation at the Hot Chips conference in Stanford. “Jalapeno can serve more AI work per unit of power, while also returning responses more quickly. It is very efficient to serve a lot of customers, but it can also be very low latency.” The InferenceX benchmark measured full end-to-end serving of an AI request, from input processing through output generation, under conditions ranging from high-throughput batch serving to interactive low-latency use. Models tested included GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1 trillion parameters, the last being the largest public model in the comparison. On GPT-OSS, Jalapeno hit approximately 1,400 tokens per second per user. On DeepSeek R1, it topped 700 tokens per second on a single concurrent request. On Kimi K2.5 1T, the chip achieved 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system, which Nvidia identifies as Vera Rubin operating at 900 to 1,150 watts. Jalapeno is rated at 700 watts, though its measured sustained power stayed at or below 550 watts on the workloads tested.

How OpenAI Built the Chip in Nine Months

Jalapeno reached tapeout in approximately nine months, a pace that semiconductor engineers have described as unprecedented for advanced ASIC silicon at the 3-nanometer node. The total development cycle from initial kickoff to fabricated silicon took about 16 months, but the core design phase occupied only nine of those months. SemiAnalysis CEO Dylan Patel, whose firm independently verified some of the benchmark runs on-site at OpenAI labs, called the results remarkable. “Usually first-generation chips are not competitive, but OpenAI is beating Nvidia Blackwell and even Rubin,” Patel wrote in his analysis. OpenAI used its own AI models throughout the development process. Earlier model generations assisted with chip design and verification, while the company latest models accelerated programming and optimization of the hardware. The chip was designed with a deliberately clear programming model, using local tensors, explicit communication, and predictable synchronization, making it tractable for both human engineers and AI coding tools. Using Codex with GPT-Astra, OpenAI brought three open-weight models that were not part of Jalapeno original production plan to high performance within two months. For selected attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-expert-written versions. “We used AI to design the chip, and designed the chip so AI could program it,” the company said in its technical blog post.

A Broader Challenge to Nvidia Dominance

The results land at a moment when Nvidia controls roughly 60 percent of TSMC CoWoS advanced packaging capacity through 2026, creating a bottleneck that has forced the entire industry to explore alternatives. OpenAI CFO Sarah Friar emphasized that Jalapeno complements existing partnerships with Nvidia, AMD, AWS, Cerebras, and CoreWeave rather than replacing them. Each of those companies is also building its own AI silicon, creating an increasingly competitive landscape. The SIA-Deloitte study released in June 2026 estimated that annual revenue from chips deployed in AI data centers could exceed 1.2 trillion dollars by 2028, a nearly tenfold increase over four years. SemiAnalysis suggested that Nvidia “CUDA moat” may not hold as firmly as once assumed, given how quickly OpenAI brought up new models on its custom silicon. Jalapeno is not a training chip and is not tuned exclusively to OpenAI own models. It is a general-purpose LLM inference accelerator, meaning it competes directly with Nvidia data center products on the efficiency metrics that hyperscalers care about most: throughput per watt and latency per dollar of compute. OpenAI plans to begin deploying Jalapeno within its own compute infrastructure by the end of 2026. Gen 2 is already deep in development, and Gen 3 is taking shape, the company said. Internal testing showed Jalapeno advantage widening further on frontier OpenAI models, suggesting the architecture becomes more valuable as workloads grow larger and more demanding.

Caveats and Market Context

There are important caveats. Nvidia and AMD have published results on larger models, including Deepseek V4 Pro and Kimi K3, that have not been tested on Jalapeno. While Vera Rubin systems are already shipping to customers, Jalapeno has reportedly not moved beyond engineering samples, with volume production expected in 2027. The benchmark comparisons also showed mixed results on total cost of ownership per token, where Jalapeno and Vera Rubin came out roughly even despite the efficiency advantage. The market implications extend beyond a single company. If OpenAI can design competitive inference silicon in nine months using its own AI tools, it reshapes the economics of the AI hardware industry. Every major cloud provider and hyperscaler is now evaluating whether custom silicon can reduce their dependency on a single vendor. Google, Amazon, and Microsoft all have active chip programs, and the CXMT surge in China memory market is adding supply chain pressure that affects everyone.

SourcesOpenAI; SemiAnalysis; TechCrunch; The Decoder; CNBC; SIA-Deloitte Report, June 2026
React to this dispatch
Share this dispatch X WhatsApp Bluesky Report an error
Written by

Founder and editor of Pulse of Nations, an independent wire service covering war, geopolitics, markets and technology.

discussion

Leave a Reply

Next dispatch Google Launches Gemini 3.5 Transcribe for Real-Time Speech AI Read →