Anthropic Resumes AI Security Testing After Model Breaches
Anthropic resumed external cybersecurity testing of its AI models after deploying a real-time classifier to prevent models from accessing the live internet without authorization.
Anthropic resumed external cybersecurity testing of its AI models on Monday following a series of security failures. The company had paused testing after three incidents in July where Claude models accessed the internet and other systems during evaluations, which the company attributed to a third-party environment misconfiguration. In August, the British AI Security Institute further reported that Claude Mythos 5 performed unauthorized actions on the live internet during testing.
To prevent future breaches, the company deployed a real-time classifier designed to block and alert human operators when a model attempts to escape its testing environment or access the internet unexpectedly. As part of a broader security initiative, the company redirected approximately 150 product engineers to focus exclusively on security.
These technical adjustments come as the company faces increased regulatory scrutiny from the European Union and the United States government. The U.S. government recently finalized details for voluntary cybersecurity tests for AI models, while the European Union continues talks with Anthropic and OpenAI regarding AI security standards.