How Zero-Day Discovery Became a Subscription Feature
In twelve months, the AI labs moved the ability to find and exploit software flaws from a contained safety risk to a subscription product — and the safety research meant to keep pace still hasn't shipped.
In September 2025, Anthropic's Frontier Red Team classified the ability to discover zero-day vulnerabilities as an AI Safety Level 3 risk — the tier that keeps a model from shipping until the danger can be contained. The team's stated purpose was to find these risks as fast as possible and tell the world about them [1]. Twelve months later, OpenAI announced GPT-6 Astra, a model that can independently find and exploit unknown flaws in real-world software, and described its rollout as a wider one to paid users and cloud providers [2]. Same capability, two registers. The distance between those two sentences was covered in four steps, each one widening who could get the tool. In March, OpenAI launched Codex Security, its first productized cyber capability, available to ChatGPT Enterprise, Business, Edu, and Pro customers. It scanned 1.2 million commits in its first month and flagged 792 critical and 10,561 high-severity issues [3]. The gate was enterprise customers. In April, both labs shipped specialized cyber models that could discover zero-days — OpenAI's GPT-5.4-Cyber and Anthropic's Mythos Preview — but restricted them to vetted partners only: Anthropic through Project Glasswing, OpenAI through a program of vetted professionals [4]. The gate was vetted partners. In June, OpenAI's Daybreak program widened it again, putting GPT-5.5-Cyber in the hands of vetted defenders — and reporting 30,000 repositories scanned and more than 500,000 findings fixed [5]. The gate was vetted defenders. In September, Astra added a new phrase to the sequence. Its advanced cyber features first go through the Daybreak Blue early-access program with select partners — Cisco, Cloudflare, Palo Alto Networks — before what OpenAI called a wider rollout to paid users and cloud providers [2]. The vetting didn't disappear; it just stopped being the whole story. For the first time, the gate opened to anyone with a subscription. The reason the line between defensive and offensive has always been thin is that the model doesn't know which side it's on. The model that finds and patches a zero-day finds and exploits it. A volunteer team made the point in August: using commercial models from OpenAI, Anthropic, and others, it surfaced 720 high- and critical-severity vulnerabilities across 390 projects in under 30 hours — roughly one critical exploit per hour per person [6]. The work was framed as defensive red-teaming. That framing is exactly the route around the guardrails that still block a direct request to attack. Meanwhile, the safety research meant to keep pace has not reached the products shipping. Anthropic's GRAM method, designed to isolate dangerous knowledge into modules that can be switched off, is by the lab's own description preliminary and not applied to production models [7]. OpenAI hired a head of preparedness in February who warned of risks of extreme and even irrecoverable harm [8]. The infrastructure is being built; the models are already out. And it's a race, not a one-off. Anthropic's Claude Fable 5.1 and Google's Gemini 3.8 Flash Cyber have both shipped their own advanced cyber models [2]. Each lab is matching the other's move. None of this means the labs dropped every control. A Chinese researcher who wanted to run autonomous attacks had to turn to the open-weight DeepSeek model, because OpenAI's and Anthropic's filters refused the requests [9]. And the defensive use is real: the New York Stock Exchange used Mythos to find and fix vulnerabilities [10], and Daybreak patched at scale. The point is not that the controls vanished. It's that the same capability serves both sides, and the gate has moved. Palo Alto Networks CEO Nikesh Arora put the catalyst plainly [11].
I’ve been trying for eight years to tell customers they’re not ready, and [Anthropic CEO Dario Amodei] did it in one event, just by launching Mythos. — Nikesh Arora
Astra ships that event to anyone with a subscription.
- 1. Anthropic Red Team Identifies High-Risk AI Capabilities
- 2. OpenAI Releases GPT-6 Astra with Critical Cyber Capabilities
- 3. OpenAI Launches Codex Security to Automate Software Vulnerability Detection
- 4. OpenAI and Anthropic Launch Specialized AI Cybersecurity Models
- 5. OpenAI Launches Daybreak Program to Automate Cyber Defense
- 6. AI Red Team Finds Hundreds of Bitcoin Exploits
- 7. Anthropic Develops GRAM Method to Isolate Dangerous AI Knowledge
- 8. OpenAI Hires Dylan Scandinaro as Head of Preparedness
- 9. Chinese Researcher Uses DeepSeek AI to Automate Cyber-Attacks
- 10. New York Stock Exchange Uses Anthropic AI to Fix Vulnerabilities
- 11. Palo Alto Networks CEO Warns of $1 Trillion Security Debt