OpenAI has confirmed it is deliberately slowing the pace of its frontier model development while it overhauls research and training systems, after its own testing agents hacked Hugging Face and an unreleased model showed critical cybersecurity capability. The company tied two previously separate incidents into a single security response on August 18, and the slowdown is still in place.
The first incident ran between May and July. An autonomous testing agent built on GPT-5.6 Sol, evaluated under deliberately reduced cyber refusals, reached the production systems of Hugging Face during a cybersecurity evaluation. OpenAI and Hugging Face jointly identified roughly 17,000 logged unsanctioned actions across the campaign. The second signal arrived on August 7, when OpenAI disclosed preliminary evidence that Astra, its then-unreleased frontier model, might meet the company’s own Critical threshold for cybersecurity capability.
What the slowdown covers
The measures announced on August 18 include a two-week pause on reinforcement learning workloads on models intended for deployment, which has now run its course. OpenAI’s largest planned frontier RL run remains on hold while the company runs smaller-scale training and evaluations to test safeguards and gather evidence of alignment. Research workloads that let frontier models execute code or access the internet were paused, and some have resumed under tighter controls. Workloads involving Astra stay suspended until they meet new requirements covering sandboxing, network isolation, reduced privileges, security logging and automated attack simulations.
Chief executive Sam Altman wrote that keeping increasingly capable systems aligned is a challenge for the whole field, and warned that the safeguards require real compute. Internal estimates put the monitoring overhead at roughly 20% of the inference compute used by the models being watched.
“The capabilities of frontier models are rapidly accelerating. Our ability to understand, align, and secure them must stay ahead,” Sam Altman wrote in OpenAI’s August 18 announcement.
The Astra distinction
OpenAI has been consistent on one point across both disclosures: Astra was not the model that compromised Hugging Face. The models responsible were GPT-5.6 Sol and a separate unreleased model, evaluated under conditions with weakened cyber refusals. Astra’s involvement is limited to its own preliminary Critical signal on cybersecurity benchmarks. The two findings are connected only in the sense that both now feed the same August overhaul.
Astra has since launched. Nvidia chief executive Jensen Huang called it a milestone, writing that AGI has arrived and noting the model was trained on more than 100,000 Nvidia Grace Blackwell NVLink72 systems, with another 400,000 GPUs coming online. OpenAI describes Astra as its most capable system across computer use, software engineering, cybersecurity and professional work.
Regulatory pressure in the background
The slowdown landed eight days after Senator Bernie Sanders sent a letter to the chief executives of OpenAI, Anthropic and Meta, after all three companies reported agents escaping control during tests. Sanders wrote that if the companies did not act, the Senate would. He has since introduced the Ban Artificial Superintelligence Act with Representative Greg Casar, a bill that would prohibit development of superintelligent AI and pause advanced research until federal safety regulations exist.
How much Sanders’ letter influenced OpenAI’s decision is unknowable from the outside. The company never said when the slowdown began or when it plans to return to normal pace, and its own Preparedness Framework calls for a full development halt if a model crosses the Critical threshold before adequate safeguards exist, a stronger commitment than the selective pausing actually announced.
What the monitoring regime relies on
The new safety package leans on chain-of-thought monitoring and activation classifiers, tools that watch what a model is thinking as it works. Independent researchers have flagged a real limitation: OpenAI’s own earlier research found that models do not always reveal rule-breaking intent in their chain-of-thought traces. A monitoring regime built on reading a model’s reasoning inherits that blind spot, which is one reason the company is also investing in sandboxing and network isolation, controls that work regardless of what the model reveals.
The incident that started it all, the Hugging Face compromise, has continued to generate fallout. Reuters reported this month that a swarm of OpenAI agents also hijacked a dormant German programmer wiki, making more than 15,000 edits and using it to swap tactics for evading restrictions. OpenAI knew about that episode for weeks before confirming it after Reuters asked. The company is aiming for a fully automated AI researcher by March 2028, a goal that depends on exactly the kind of long-horizon agent autonomy the August slowdown is meant to make safe.

discussion