OpenAI Warns Next-Gen AI Models Pose High Cybersecurity Risk
OpenAI announced a defense-in-depth strategy and a new risk council to prevent its increasingly powerful AI models from facilitating zero-day exploits and industrial intrusions.
OpenAI warned that its next-generation artificial intelligence models may reach high cybersecurity risk levels, potentially enabling the development of zero-day remote exploits and complex industrial intrusion operations. This warning follows a significant increase in model performance on capture-the-flag challenges, which rose from 27 percent on GPT-5 in August to 76 percent on GPT-5.1-Codex-Max in November. Researcher Fouad Matin identified the models' ability to operate for extended periods as a primary driver of these risks.
To mitigate these threats, the company is implementing a defense-in-depth strategy including access controls, infrastructure hardening, and the private beta of Aardvark, an agentic security tool designed to identify vulnerabilities and suggest patches. OpenAI also established the Frontier Risk Council, an advisory group of security practitioners, and a trusted access program to provide cyberdefense users with enhanced model capabilities.
Industry reactions to these measures are mixed. Some experts, including representatives from Recorded Future and ThreatAware, suggest that AI-driven threats do not exceed the capabilities of organizations following basic security fundamentals. Other critics argue that safety frameworks and refusal training only constrain non-malicious users and will not stop determined adversaries, who may instead turn to open-weight models that lack safety controls.