Skip to content
live markets
S&P 5007,798.99▲ 3.77%NASDAQ26,803.03▲ 3.59%DOW53,839.99▲ 2.56%GOLD4,400.90▲ 8.37%WTI82.50▲ 3.98%BRENT88.00▲ 3.86%EUR/USD1.1558▲ 1.53%USD/JPY159.13▼ 2.03%DXY99.74▼ 1.19%BTC$62,795▼ 1.50%ETH$1,873▼ 1.20%SOL$75.68▼ 1.10%TOTAL CRYPTO$2.25T▼ 1.09%
pulseofnations.
Fri, Aug 14 2026 — 08:50 UTC telegram ↗ bluesky ↗ Join the wire

OpenAI Launches Ultrafast Mode, Running GPT-5.6 Sol at 14x Speed

OpenAI introduces Ultrafast, a new service tier powered by Cerebras that runs GPT-5.6 Sol at up to 750 output tokens per second, targeting enterprise real-time workflows.

OpenAI on Thursday unveiled Ultrafast, a new service tier that runs its flagship GPT-5.6 Sol model at up to 14 times the speed of standard processing, delivering as many as 750 output tokens per second.

The mode is powered by OpenAI’s partnership with chipmaker Cerebras, whose wafer-scale processors are designed for extreme throughput. “Until now, getting real-time speed typically meant choosing a smaller or more specialized model,” OpenAI said in its announcement. “Ultrafast points to progress in a new direction: more useful work per second.”

GPT-5.6 Sol, released in July, is OpenAI’s most capable reasoning model with a 1.05 million token context window. Standard processing delivers roughly 62 tokens per second. Ultrafast pushes that to 750 tokens per second, a leap that matters most for agentic workflows where dozens of sequential model calls can take minutes at normal speed.

OpenAI envisions the mode being deployed in incident response, customer service and support, financial market analysis, and e-commerce operations where every second of latency has a direct business cost. The company said a 40-step agent workflow that spent eight seconds per step at standard speed now completes closer to two seconds per step on Ultrafast.

The announcement intensifies the speed race among frontier AI labs. Anthropic offers a fast mode for Claude, though it does not approach the throughput OpenAI is claiming. Google has also invested in accelerated serving infrastructure for Gemini, and Nvidia’s GPU ecosystem remains the default for most production AI workloads.

Ultrafast is launching first as a limited preview available to select API customers. OpenAI said it plans to expand access as Cerebras scales capacity. The tier is initially API-only, meaning it is aimed at developers building production applications rather than individual ChatGPT users.

The move underscores OpenAI’s strategy of differentiating on infrastructure as much as model quality. By partnering with Cerebras rather than relying solely on Nvidia’s GPU clusters, OpenAI gains a second supply line for compute at a time when demand for frontier AI serving far outstrips available capacity across the industry.

Sources: OpenAI Blog, TechCrunch, Fierce Network

React to this dispatch
Share this dispatch Telegram X WhatsApp Report an error

discussion

Join the discussion

Your email address will not be published. Required fields are marked *

Next dispatch Anthropic Watermarks Claude Outputs Amid User Backlash Read →