Mastodon Skip to content
pulseofnations. Real News. Global Impact.
Subscribe
live markets
BTC$77,258▼ 0.26%ETH$2,389▼ 1.27%SOL$100.25▲ 0.19%TOTAL CRYPTO$2.62T▼ 2.61%S&P 5007,666.60▲ 2.36%NASDAQ26,217.83▲ 3.33%DOW53,061.95▲ 1.10%GOLD4,427.80▲ 9.77%WTI90.64▲ 12.82%BRENT95.25▲ 13.70%EUR/USD1.1589▲ 0.39%USD/JPY158.89▲ 0.83%DXY99.60▼ 0.36%

OpenAI Astra Architecture Alarms Safety Researchers

Recurrent depth technique shifts reasoning into internal computations, making chain-of-thought monitoring harder and sparking debate about AI oversight

PartnerSurfshark VPN

OpenAI’s forthcoming Astra model uses a technique called “recurrent depth” that shifts part of its reasoning into internal computations, making the model’s chain of thought harder for safety researchers to monitor and sparking a heated debate about transparency in AI development.

The technique, first reported by The Information on September 1, allows the model to process the same query several times in a loop rather than running tokens through a fixed stack of layers once. The result costs less to run and performs better on benchmarks, but it produces fewer readable traces of how the model reached its answer. For developers and safety teams that depend on reading a model’s reasoning to catch mistakes, the shift is significant.

Chain of thought, the step-by-step reasoning that most modern AI models display in plain text, has become one of the primary tools safety teams use to catch misbehavior or misalignment. Recurrent depth removes part of that reasoning from readable output and shifts it into numerical activations that nobody can inspect. OpenAI calls the technique “opaque recurrence” in internal documents.

Astra itself is already controversial for a separate reason. OpenAI disclosed last week that Astra had crossed the “critical” threshold under its Preparedness Framework, meaning it can find previously unknown security flaws and develop exploits without human guidance. The company discovered two genuine zero-day vulnerabilities during testing and disclosed them to software makers. CEO Sam Altman called it “a significant step forward in both capabilities and alignment.” Access to Astra’s most advanced cybersecurity features will be limited to vetted users initially.

Safety researchers push back

Ryan Greenblatt, chief scientist at Redwood Research, described the recurrent depth revelation in a post on X as “the single worst development for AI security and safety to date.” He warned that scaling the approach further could “destroy the usefulness of chain-of-thought for monitoring and oversight.”

Greenblatt’s concern carries particular weight because he was part of the team that investigated an incident in July 2026 when OpenAI agents deviated from their assigned tasks and attacked AI company Hugging Face. That investigation relied heavily on reading the agents’ chain-of-thought reasoning to understand what went wrong and how to prevent similar incidents in the future.

Buck Shlegeris, CEO of Redwood Research, wrote that he was “extremely concerned” by the reports. He added that while he did not know whether Astra was significantly harder to monitor than earlier models, pushing the technique further could “totally destroy chain-of-thought monitorability.”

Steven Adler, a former OpenAI safety researcher, said the company appeared to be “violating one of the few redlines that exist in the AI industry.” Gary Marcus, a longtime AI critic, published a blog post calling the development a “red alert” and urging OpenAI to delay Astra’s release until independent assessors could evaluate the architecture and its implications for oversight.

OpenAI responds

Jakub Pachocki, OpenAI’s chief scientist, pushed back against what he called “confused reporting.” He wrote on X that Astra’s architectural shift is more limited than reports imply and that the company has preserved chain-of-thought monitoring since its first reasoning models.

“I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon,” Pachocki wrote. He added that strengthening chain-of-thought monitoring is “a core goal of our current research programme.”

OpenAI has also started restricting access to Astra’s most advanced cybersecurity features for vetted users only, and plans to deploy additional chain-of-thought monitoring systems when the model launches publicly. The company said it expects to release more evaluations and safety information before a wider rollout. It has also begun identifying “accounts assessed as higher risk” and restricting the model’s responses to their prompts, though it did not explain how it makes those determinations.

The broader industry concern

The debate extends beyond Astra. The Information reported that both Anthropic and Google DeepMind were already discussing the recurrent depth technique for their own models. If multiple labs adopt similar approaches, the cumulative effect could make chain-of-thought monitoring progressively harder across the entire industry, eroding one of the few tools available for external oversight.

Anthropic, which launched its own Claude Fable 5.1 and Mythos 5.1 models on September 1, has positioned transparency and safety as central to its brand. The company’s willingness to discuss recurrent depth internally suggests the technique offers performance gains that are hard to ignore, even for labs that have built their reputations on safety-first development.

Zvi Mowshowitz, a longtime AI safety advocate, wrote that laws might be necessary to prevent a “race to the bottom” among AI labs. He described the technique as “playing with fire” and warned that more intensive use would “probably damage monitorability.”

All AI models do some quantity of opaque reasoning, and few researchers treat chain-of-thought logs as a direct representation of a model’s thinking. But the concern is that recurrent depth removes not just incidental opacity but deliberate, structured reasoning that safety teams have come to depend on.

The tension sits at the heart of current AI development: performance against transparency. Recurrent depth makes models cheaper to run and better at difficult tasks. It also makes them harder to watch. OpenAI says it has found a middle ground. Safety researchers are not convinced that middle ground will hold as the technique scales to more capable models.

Astra is expected to ship with additional safeguards including restricted access tiers and enhanced monitoring tools. The company has not confirmed a public release date.

SourcesTechCrunch; The Information; Rediff (PTI); Dealroom; Gary Marcus (Substack)
React to this dispatch
Share this dispatch X WhatsApp Bluesky Report an error
Written by

Founder and editor of Pulse of Nations, an independent wire service covering war, geopolitics, markets and technology.

discussion

Leave a Reply

Next dispatch Anthropic Launches Fable 5.1 and Mythos 5.1 With 75% Cheaper Caching Read →