Hot Chips 2026, the semiconductor industry’s premier technical conference, opened Monday at Stanford University with NVIDIA, AMD and Intel all presenting next-generation GPU architectures in a single day, underscoring the intensity of the race to build silicon for agentic AI workloads.
The sold-out event features presentations from the architects behind the chips themselves, making it the most detailed public disclosure of AI silicon roadmaps each year. NVIDIA’s Manas Mandal and Rajballav Dash are presenting the Rubin GPU under the session title ‘Driving the Era of Agentic AI,’ while AMD’s Alan Smith and Maiyuran Subramaniam detail the Instinct MI400 Series. Intel closes the GPU session with Crescent Island, described as a GPU designed specifically for AI inference.
NVIDIA Rubin: 336 Billion Transistors
NVIDIA’s Rubin GPU, designated R200, is built on TSMC’s 3nm process using a multi-chip module design with two compute dies and two I/O dies. The chip carries 336 billion transistors, 1.6 times the count of the current Blackwell generation, and delivers 50 PFLOPS of FP4 inference performance. It pairs with 288 GB of HBM4 memory across eight stacks, providing 22 TB/s of bandwidth, a 2.75x improvement over Blackwell.
The full Vera Rubin platform includes the Vera CPU, a next-generation DPU, NVLink 6 networking and Ethernet switching. NVIDIA claims the system delivers a 10x reduction in inference token cost versus Blackwell for Mixture of Experts models, though that benchmark is specific to the Kimi-K2-Thinking model at particular sequence lengths.
AMD MI400: The Open-Software Challenger
AMD’s Instinct MI400 series, built on the CDNA 5 architecture, targets both AI training and high-performance computing. The company has secured partnerships with Meta, Microsoft Azure and European sovereign AI programs. AMD is positioning the MI400 on total cost of ownership and memory capacity advantages, particularly for inference workloads where memory bandwidth is the bottleneck.
AMD also presented the system architecture of the MI400 in a separate session, signaling the company’s push to compete not just at the chip level but at the rack scale where NVIDIA has dominated.
Intel Crescent Island for Inference
Intel’s Crescent Island represents a different strategic bet, focusing exclusively on AI inference rather than training. The GPU is designed to optimize cost-per-token for deployed models, a metric that becomes increasingly important as AI applications move from development to production at scale.
The convergence of three major GPU reveals at a single event is unusual and reflects how the AI chip market has moved beyond NVIDIA’s solo dominance. With Gartner forecasting semiconductor revenue to hit $1.6 trillion in 2026, driven largely by AI infrastructure, the competition among chipmakers has never been more commercially significant.
Waymo’s Daniel Rosenband is also delivering the day’s keynote, titled ‘Compute in Motion: Challenges of Autonomous Driving,’ revealing details about the company’s custom 5nm sensor fusion ASIC built on TSMC’s automotive-grade N5A process. OpenAI’s Richard Ho is presenting ‘You Can Just Build Chips,’ signaling the AI lab’s own silicon ambitions.
discussion