OpenAI Announces Astra Model With Critical Cyber Capabilities
OpenAI is releasing Astra, an AI model capable of independently finding and exploiting software vulnerabilities, with advanced features restricted to select security partners.
OpenAI announced the upcoming release of Astra, its first AI model to reach a critical cyber capability threshold. The company defines this threshold as the ability to independently find and exploit previously unknown vulnerabilities in real-world software, including the capacity to chain multiple exploits to penetrate target systems.
To manage the risks associated with these capabilities, the company implemented a multi-week development pause to establish safety controls. These measures include a misalignment monitor designed to resist jailbreaking and refuse unsafe queries. Astra demonstrated high performance in testing, scoring 100 percent on the ExploitBench benchmark and outperforming models such as GPT-5.6 Sol and Anthropic's Mythos.
While a general version of Astra will be released soon, OpenAI is restricting advanced cyber capabilities to a select group of partners through the Daybreak Blue early-access program. This group includes Cisco, Cloudflare, and Palo Alto Networks, who will use the model to help harden digital defenses. The announcement follows a broader industry trend toward security hardening, with Anthropic pausing training workloads and Meta disclosing incidents involving the cybersecurity capabilities of its own models.