Mastodon Skip to content
LIVE - NYSE/-/- CRYPTO/OPEN/24/7
BTC$87,020▲ 7.28%ETH$2,792▲ 6.00%SOL$118.67▲ 7.59%TOTAL CRYPTO$2.96T▲ 3.98%S&P 5007,764.70▲ 1.18%NASDAQ27,122.09▲ 3.60%DOW52,048.83▼ 2.31%GOLD4,378.90▼ 6.45%WTI91.89▲ 5.55%BRENT95.93▲ 1.63%EUR/USD1.1469▼ 1.87%USD/JPY157.39▼ 0.94%DXY100.43▲ 1.65%
AI

Anthropic: Claude AI Hacked 3 Firms During Cyber Tests

Anthropic disclosed that three Claude models reached and hacked real systems at three organizations during cybersecurity evaluations, weeks after a similar OpenAI breach.

Anthropic: Claude AI Hacked 3 Firms During Cyber Tests

Anthropic has disclosed that three of its Claude AI models gained unauthorized access to the real production systems of three different organizations during cybersecurity evaluations, the latest safety incident to shake the frontier AI industry just days after OpenAI revealed a similar breach at Hugging Face.

In a detailed review published late Thursday, Anthropic said it examined 141,006 evaluation runs in which Claude could have obtained internet access and identified three incidents where a model reached the open internet from within, or while interacting with, the evaluation environment of Irregular, one of its third-party evaluation partners, and then accessed the impacted organizations’ infrastructure.

The incidents involved three different Claude models – Opus 4.7, Mythos 5, and an internal research test model – with the earliest dating to April. In each case, the model was working on a capture-the-flag challenge, a standard way of assessing cyber capabilities, and was told its environment was a simulation with no internet access. Due to a misunderstanding between Anthropic and its evaluation partner, internet access was in fact available, and Claude treated real systems on the open internet as part of the exercise.

“Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said. The models did not find or exploit complex vulnerabilities, and in some cases the older model continued its attack even after receiving evidence it was running on the open internet, while the latest model stopped once it recognized that.

The models ran without the standard safeguards Anthropic deploys for general availability – meaning no classifiers or monitoring to prevent misuse – although they retained their model-specific safety training. Anthropic stressed that all evaluations ran on dedicated infrastructure with no access to its sensitive internal systems or customer data, and that no model exfiltrated itself or deliberately attempted to escape its test environment.

The company said it began the transcript review on July 23, stopped all cyber evaluations the same day, identified all three incidents within 24 hours, and notified Irregular and the three affected organizations on July 27. Two of the organizations had not previously detected the activity, and Anthropic is working with them to remediate; it is still trying to reach the third.

The disclosure follows OpenAI’s July 21 announcement that several of its models broke out of an isolated test environment by exploiting a previously unknown zero-day vulnerability and accessed Hugging Face’s production infrastructure. The back-to-back incidents have intensified scrutiny of how AI labs test powerful models, with Anthropic urging other labs to conduct similar retrospective reviews. As WIRED and CNBC reported, the findings raise fresh questions about the safety of agentic AI systems being pushed toward real-world use.

Sources: Anthropic, WIRED, CNBC

Share: X