# Frontier AI Lab Reports Model Crossed Threshold on Dangerous-Capability Evaluation

A leading AI lab disclosed that its newest frontier model crossed an internal danger threshold on a cybersecurity-uplift evaluation, automatically triggering restricted release while additional safeguards are built.

A leading AI research lab has disclosed that its most capable model to date crossed an internal threshold on one of its dangerous-capability evaluations, and that the result, not a policy decision, is what triggered restricted release.

## What these evaluations actually test
Frontier labs increasingly run a standard battery of tests before release, checking whether a new model provides meaningful uplift for cyberattacks, offers hazardous-material guidance beyond what's already public, or shows autonomous planning ability serious enough to warrant closer oversight. The evaluations use a rubric agreed on in advance, not an ad hoc judgment call made after the fact.

## What actually happened here
According to the lab, the model exceeded its threshold specifically on the cybersecurity[↗](/cybersecurity)-uplift evaluation. That doesn't mean the model behaves unsafely by default, evaluators found that with sufficiently persistent, adversarial prompting, it could provide a level of offensive-security assistance the lab's own policy doesn't allow for unrestricted release.

## So what happens to the model now?
It stays in restricted internal testing while the lab develops mitigations, which could mean stronger refusal training on the specific evaluated tasks, added runtime monitoring, or a narrower release, vetted enterprise access, say, rather than public availability.

Disclosures like this are becoming more common as evaluation frameworks mature industry-wide. Read as intended, this one is a success story: a concerning capability caught before broad release, not after.

Source: [Anthropic: Activating AI Safety Level 3 Protections](https://www.anthropic.com/news/activating-asl3-protections)
