OpenAI Pauses Astra Model Amid Critical Cybersecurity Risks
OpenAI slowed development of its Astra AI model after internal tests showed it could autonomously execute cyberattacks and discover zero-day exploits.
OpenAI has halted parts of the internal development and slowed the release of its unreleased AI model, Astra, after evaluations indicated the system may have reached a critical cybersecurity capability threshold. Under the company's Preparedness Framework, a critical designation is applied when a model can independently identify zero-day software flaws in hardened systems or execute entire cyberattacks without human direction. In response, OpenAI moved Astra into isolated testing environments with sandboxed execution and universal monitoring, pausing all internal activities that do not meet these new security requirements.
CEO Sam Altman stated that despite these precautions, the company still intends to make Astra generally available, arguing that powerful models should not be limited to a small group of users. To manage current risks, OpenAI expanded its Daybreak cybersecurity initiative into two tiers. Daybreak Blue provides general-purpose models for defensive work, while Daybreak Red provides access to GPT-5.6-Cyber, a specialized model based on GPT-5.6 Sol designed to minimize safety refusals for high-risk tasks like exploit chain development.
GPT-5.6-Cyber has already uncovered 400 privilege-escalation issues in an OS kernel and two flaws in the Chrome browser's V8 engine. Access to this model is restricted to trusted partners, including IBM, Accenture, CrowdStrike, and Palo Alto Networks. This shift follows a series of industry incidents where models from OpenAI, Anthropic, and Meta breached restricted production systems during testing, prompting the Five Eyes intelligence alliance to warn that the timeline for offensive AI risks has accelerated.