Anthropic has disclosed the existence of an unreleased internal model it calls “Model 2,” which is notably more capable than its current flagship Claude Mythos 5.
The details emerged in the company’s latest AI alignment report, published on Thursday. Anthropic stated that it has no current plans to release the model externally, as it has not yet completed full pre-deployment safety assessments.
The report also marked a shift in the company’s safety outlook, raising its rating for catastrophic misalignment risk from “very low” to “low.” Anthropic attributed the change to increased overall uncertainty and recent cybersecurity incidents involving its models, rather than a specific failed safety test.
“We are less confident in this assessment than before because our best internal benchmarks struggle to keep up with LLM advances,” the report noted regarding recursive self-improvement thresholds.
A More Capable Internal Model
Model 2 represents a significant leap in Anthropic’s internal research capabilities. While Mythos 5 is currently the most advanced model available to developers through the Claude API, Model 2 reportedly excels in reasoning and complex task completion.
Anthropic emphasized that the model is currently “undergoing rigorous testing” and will only be considered for public release once it meets their strict safety criteria. The decision to keep it internal reflects the company’s “scaling” approach to safety, ensuring that capabilities do not outpace alignment techniques.
Addressing Emerging Risks
The shift in risk rating is notable for an industry that has often downplayed the possibility of models acting against their intended goals. Anthropic cited several “real-world incidents” where models were used in ways that highlighted the need for better robustness.
The company is also proposing a new industry-wide framework for scoring “jailbreak severity,” a move aimed at creating standardized metrics for how easily a model’s safety guardrails can be bypassed. This framework is being developed in collaboration with Amazon, Microsoft, and Google.
Sources: SiliconANGLE; Anthropic AI Alignment Report; AI Weekly
discussion