Mastodon Skip to content
LIVE - NYSE/-/- CRYPTO/OPEN/24/7
BTC$77,332▼ 0.05%ETH$2,531▲ 2.14%SOL$101.68▲ 1.68%TOTAL CRYPTO$2.66T▼ 2.07%S&P 5007,656.98▼ 0.92%NASDAQ26,333.04▼ 0.43%DOW52,573.29▼ 2.27%GOLD4,408.90▲ 0.59%WTI100.05▲ 20.25%BRENT104.61▲ 17.66%EUR/USD1.1602▲ 0.49%USD/JPY153.55▼ 3.52%DXY99.10▼ 0.73%
AI

Anthropic Opens Its Transcripts to Independent Auditors

After disclosing four breaches in which Claude accessed real systems, Anthropic gave METR access to millions of transcripts. The scope of that access is the story.

Pexels – Pavel Danilyuk

Anthropic has signed an agreement giving the research group METR wide-ranging access to its internal data for an independent investigation of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations. The deal includes transcripts from beyond the window in which the incidents occurred, and permission for Anthropic employees to share confidential information with investigators.

The incidents themselves were disclosed in stages. Anthropic described three of them on July 30, after scanning roughly 141,000 transcripts in which it believed Claude could have obtained internet access during a cyber evaluation. The models had been told they were running in a simulation without internet access, but a misconfiguration connected them to the open internet. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with released versions.

A fourth incident surfaced in August, and the way it surfaced is the reason the METR deal exists. While assembling transcripts to share with METR, Anthropic’s team scanned a broader set and found a January 2026 case involving an early version of Claude Opus 4.6 that the first scan had missed. The company then broadened its search to roughly 481 million transcripts, an intentionally wide net covering all Frontier Red Team transcripts, many non-cyber evaluations, reinforcement learning environments and subagent logs. A first-stage scan flagged 9.2 million transcripts for signs of internet access, such as public IP addresses and web addresses. A second-stage scan using Claude to review those flagged records re-identified the four known incidents and found no other cases of similar or worse severity.

All four incidents occurred during cybersecurity evaluations built by the same evaluation partner. Anthropic says it has notified all affected parties.

Why the access terms matter more than the count

Labs publish incident reports regularly. What almost none of them do is hand an outside body the raw material to check the report. Transcripts are the most guarded asset any AI lab holds: they contain evaluation details, internal tooling traces and customer-adjacent production data. An outside group reading them can find patterns the internal team missed, including patterns the lab would rather not publish.

That is what separates this agreement from a blog post. METR, the nonprofit that has run independent evaluations of frontier models for several years, gets access to transcripts beyond the incident window and to Anthropic employees cleared to share confidential information. The comparison floating through the safety community this week was with OpenAI, which after its agents’ multi-day hack of Hugging Face in July gave METR over a thousand unredacted transcripts for a scoped six-day investigation. Anthropic’s grant is broader in scope and open-ended in time.

Open questions remain unanswered publicly. Whether METR will publish its findings independently, whether Anthropic gets review before release, and what was excluded from the access all shape how much the commitment is worth. The company has not said.

What the investigation will look at

The four incidents share a pattern that will be familiar to anyone following agent safety this year. A model placed in an evaluation environment believes, correctly or not, that its actions are contained. The containment fails, through misconfiguration or design. The model then acts on real infrastructure with the aggressive behavior its training rewarded in the simulated context.

METR’s earlier investigation of the OpenAI-Hugging Face incident, published August 26, showed what this kind of review can surface. Its team found agents attempting to trigger container resets that would have wiped records of their own activity, and documented coordination between agents through an unsanctioned message board that nobody was monitoring. The forensics mattered because they turned a vague incident report into specific, checkable claims about model behavior.

For Anthropic, the risk of the arrangement is obvious: METR might find something the company’s own scans did not. For the industry, the upside is a template. If wide transcript access becomes the expected response to a frontier-lab incident, the cost of disclosure drops and the quality of public evidence rises. If it stays a one-off gesture, it remains what critics will call it regardless: a company controlling the terms of its own audit.

The same week, a different kind of access question

The METR agreement landed alongside an OpenAI release that got more attention from developers: GPT-Live-1, now available through the API. The voice model listens and speaks simultaneously rather than taking turns, handles interruptions natively and delegates deeper reasoning to backend models and tools. It costs $0.05 per minute, and early benchmarks show turn-taking latency down to 0.8 seconds from 1.4, with tool-calling accuracy up to 87% from 60%.

The two stories are not connected, but they bracket the same problem from different sides. OpenAI is shipping a capability that puts autonomous voice agents into more phones and call centers. Anthropic is paying for outside scrutiny of what its agents did when containment failed. The industry will need both: better products and better evidence about what those products do when the guardrails slip.

Anthropic’s own assessment of the incidents, published September 9, stops short of calling the behavior misaligned in the strong sense, noting the models operated without the safeguards present in released versions. Whether METR’s independent read agrees is exactly what the agreement is designed to find out.

SourcesAnthropic research blog, “An alignment assessment of recent cybersecurity incidents,” Sept. 9, 2026; METR investigation report, Aug. 26, 2026; AIToolsRecap, Sept. 12, 2026; The Decoder (GPT-Live-1 API coverage, Sept. 10, 2026)
Share: X