Mastodon Skip to content Breaking Ledger Drains $86M: Supply Chain Nightmare•Ledger Drains $86M: Supply Chain Nightmare•Ledger Drains $86M: Supply Chain Nightmare•Ledger Drains $86M: Supply Chain Nightmare•Ledger Drains $86M: Supply Chain Nightmare•
LIVE - NYSE/-/- CRYPTO/OPEN/24/7
BTC$82,552▲ 0.91%ETH$2,487▲ 0.35%SOL$109.11▼ 0.97%TOTAL CRYPTO$2.8T▼ 2.08%S&P 5007,811.54▲ 0.59%NASDAQ27,366.17▲ 0.64%DOW51,654.95▲ 0.83%GOLD4,220.30▲ 1.52%WTI91.66▲ 0.19%BRENT104.43▲ 0.14%EUR/USD1.1206▼ 0.06%USD/JPY158.25▲ 0.12%DXY102.23▲ 0.09%
AI

OpenAI Shows Off Jalapeño, Its First In-House AI Chip

Built with Broadcom and Celestica, Jalapeño is an inference accelerator OpenAI says beats Nvidia Blackwell on work per watt. Deployment starts by year-end.

Pexels – Andrew Neel

OpenAI has shown Jalapeño, its first custom inference accelerator, claiming 1.5x to 1.9x more AI work per watt than Nvidia’s best commercial systems at peak throughput across three tested models. The chip is the product of a partnership with Broadcom and Celestica, tied to an agreement to build 10 gigawatts of custom silicon, and OpenAI plans to deploy it inside its own infrastructure by the end of the year. There is no external product attached. The numbers come from a Hot Chips presentation. SemiAnalysis, the research shop that reviewed the claims, argues the fairer comparison would be Nvidia’s newer Vera Rubin platform rather than Blackwell. It also says Jalapeño still wins on tokens per megawatt even against that newer comparison. The distinction matters because Nvidia does not stand still. Blackwell launched two years ago and is now the incumbent rather than the frontier.

Inference only

Jalapeño serves models, it does not train them. That split reflects where the industry’s real compute burden now sits. Training consumes enormous single bursts of compute at a handful of facilities. Inference runs constantly across every customer request, and each serving dollar matters more to unit economics. Google trains on its own TPUs. Amazon and Meta run their own inference silicon. With Jalapeño, OpenAI stops being the exception. The company had been the only major model developer buying virtually all of its compute from Nvidia, a dependency that showed up in every margin conversation and every capital raise. Building a serving chip is the first structural answer to that dependency.

The Broadcom deal

The chip exists because of a Broadcom partnership announced in October 2025, tied to a 10-gigawatt deployment target for custom accelerators. Broadcom’s role is silicon design, and the two companies have been working on the program long enough that the Hot Chips disclosure looks less like a debut and more like a status report on hardware already running in test clusters. Schwarzmuller Celestica’s part is systems integration: assembly, cabling, power delivery and the rack-level engineering that determines whether a chip hits its thermal limits in production or not. Inference accelerators run in dense configurations where cooling failures translate directly into lost serving capacity. Also running in parallel, Broadcom has been talking to lenders about raising more than $50 billion to finance the buildout for its OpenAI work, a sum that would rank among the largest single supply-chain financings on record.

SemiAnalysis argues the fairer comparison is Nvidia’s Vera Rubin platform rather than Blackwell, and says Jalapeño still wins on tokens per megawatt against that newer system.

Metric Jalapeño vs. Blackwell
Work per watt, peak throughput, 3 tested models 1.5x to 1.9x
Throughput, interactive workloads 2.1x to 4.1x
Target use case Inference only
Deployment OpenAI internal, by year-end
External product None announced
Program scale 10 GW of custom silicon

What it does not do

Jalapeño does not train models, and it does not sell to anyone else. That second point is what makes the announcement less threatening to Nvidia than the headline numbers imply. A chip deployed inside one customer’s infrastructure does not take share from Nvidia’s commercial business. Nvidia’s data center revenue runs into the hundreds of billions annually, and its relationship with OpenAI spans training partnerships, supply commitments and equity stakes. What it does change is negotiating leverage. OpenAI can now point at a credible internal alternative when it argues about compute pricing, and that argument just got harder for Nvidia to dismiss. There is second-order risk for Nvidia here worth flagging. Every major cloud now has its own inference silicon. If each of the biggest AI consumers builds a serving chip, Nvidia’s growth in that segment depends increasingly on the mid-tier cloud providers and enterprise customers who never build their own. That is still an enormous market, but it is a narrower base than the one Nvidia had when it sold to essentially everyone.

Context in the broader buildout

The Jalapeño story fits a pattern of compute financing that has reshaped AI infrastructure in eighteen months. Oracle has been negotiating with Apollo and Goldman Sachs for large chip financing. SpaceX has raised in the tens of billions. Broadcom itself is hunting more than $60 billion in debt for Anthropic’s chip purchases. Nvidia has committed $105 billion to OpenAI’s Ohio campus, taking equity in the power company rather than only backing its debt. AMD and Anthropic announced a similar partnership with up to two gigawatts of Instinct GPUs, along with an AMD investment of up to $5 billion. IBM signed a separate multi-year $240 million deal with Together AI for an inference cluster on IBM Cloud. Each of these is a line item in a buildout that has moved from cash spend to structured finance. A single-gigawatt kernel of that structure, Jalapeño, is the first in-house piece OpenAI owns outright.

What to watch

Three signals matter. First, deployment pace. OpenAI has a year-end target for putting Jalapeño into its own serving fleet, and slippage there would be the first real datapoint against the program. Second, throughput at scale. Vendor presentations and cluster benchmarks differ on everything from thermal throttling to software maturity, and running the chip at datacenter scale is where those gaps show up and get fixed or linger. Third, whether OpenAI passes any of the savings through. The company has infrastructure savings in mind when it prices products, and mid-tier model prices have already fallen twice in recent months. Cheaper inference is also a political question. Data center power is now a public issue rather than a private one, and a chip that does more work per watt reduces power draw per unit of compute, a metric that matters in local permitting fights around every new facility. The one comparison that has not been done yet is the one that counts: serving OpenAI’s actual user traffic at actual cost per million tokens, for a full quarter, with production thermal and failure rates included. Benchmarks from a conference stage are designed specifications under ideal conditions. Production reality has regressed from those numbers in every custom silicon program the industry has seen so far. The year-end deployment date gives the market a concrete line to watch.

SourcesCNBC; SemiAnalysis; the-decoder.com; Broadcom.
Share: X