ThinkPatternGet the app
Story
TECHNOLOGY · SEP 4, 2025

Anthropic Red Team Identifies High-Risk AI Capabilities

Anthropic uses its Frontier Red Team to stress-test AI models for cybersecurity and nuclear risks to inform its responsible scaling policies.

The Frontier Red Team, a group of 15 researchers at Anthropic, is stress-testing the company's AI models to identify critical risks in cybersecurity, biological research, and national security. During the 33rd annual DEF CON in Las Vegas, researcher Keane Lucas demonstrated that the Claude model family can outperform humans in simulated hacking contests, illustrating how criminal or state actors could potentially weaponize the technology.

These findings directly influence Anthropic's responsible scaling policy. The company released Claude Opus 4 under AI Safety Level 3 because the model demonstrated an ability to assist in the production of chemical, biological, radiological, or nuclear weapons. To further mitigate these threats, Anthropic partnered with the National Nuclear Security Administration, an agency within the U.S. Department of Energy, to develop tools that flag dangerous nuclear-related conversations.

While Anthropic presents these efforts as a commitment to safety, the approach has drawn criticism. Nvidia CEO Jensen Huang suggested the company's transparency is a strategy for regulatory capture, intended to create rules that benefit Anthropic over its competitors.


Reported across 3 outlets
Actors
AnthropicLogan GrahamJack ClarkNational Nuclear Security Administration

Keep reading in the app

The full story and every source, free in the app.

Download on the App StoreComing soonGoogle Play