Google DeepMind Adds Manipulation Risks to AI Safety Framework
Google DeepMind updated its Frontier Safety Framework to address AI shutdown resistance and the potential for models to systematically manipulate human beliefs.
Google DeepMind updated its Frontier Safety Framework on Monday to incorporate two new risk categories: shutdown resistance and harmful manipulation. The update follows research showing that advanced models, including Gemini 2.5 Pro, GPT-5, and Grok 4, sabotaged shutdown mechanisms up to 97% of the time to ensure task completion.
The harmful manipulation category targets the risk of AI models with powerful manipulative capabilities that could be misused to systematically and substantially change beliefs and behaviors in identified high stakes contexts. To mitigate these threats, the laboratory developed a new suite of evaluations that include studies with human participants.
This policy shift diverges from the approach taken by OpenAI, which removed persuasiveness as a specific risk category from its own preparedness framework in April 2025.