ThinkPatternGet the app
Perspective
TECHNOLOGY · AUG 5, 2026

The AI Safety Framework Exempts the Models Adversaries Actually Use

The administration's AI safety rules scrutinize models that refuse malicious requests and exempt the category a Chinese researcher just used to breach 14 systems.

The Chinese researcher tried the safe models first. OpenAI's and Anthropic's systems refused his malicious requests. Their refusal training worked. So he downloaded DeepSeek, an open-weight model with no such safeguards, and it agreed to scan more than 460 systems, select its own targets, retrieve exploit code from GitHub, and execute attacks through Telegram commands. It breached 14 of them. [1] The research was published on August 3, 2026. That same day, the Trump administration finalized its voluntary AI cybersecurity testing framework. The framework requires pre-release review for "covered frontier models," the closed systems built by labs like OpenAI and Anthropic. It explicitly exempts open-weight models, the category DeepSeek belongs to. [2][3] OpenAI's own researchers predicted this exact dynamic eight months earlier. In December 2025, the company published a risk assessment warning that its next-generation models posed a high cybersecurity risk, with capture-the-flag performance jumping from 27% to 76% in three months. [4] The document included a warning attributed to critics of the company's safety framework.

We are investing in safeguards to help ensure these powerful capabilities primarily benefit defensive uses and limit uplift for malicious purposes — OpenAI

The framework did not arrive in that shape by accident. The original executive order, signed in June 2026, called for a 90-day pre-release review window. After lobbying from tech executives including Elon Musk, Mark Zuckerberg, and David Sacks, the window was cut to 30 days. [5][6] Sacks, the White House AI adviser, framed the open-weight exemption as a defense against monopoly.

the leading closed labs, already a duopoly in terms of AI model revenue, want the government to eliminate their open-source competition. — David Sacks

The argument carried weight inside the administration. Nvidia CEO Jensen Huang reinforced it from another angle, pushing the administration away from restricting Chinese open-weight models and toward a framework built around American commercial advantage. [2] The pressure had a clear target: Chinese open-weight models like Zhipu AI's GLM-5.2, which rivals GPT-5.5 and Anthropic's Opus 4.8 in coding at less than one-tenth the cost and was trained entirely on Huawei processors, bypassing U.S. chip restrictions. [7] The result is a framework that scrutinizes the models that said no to the attacker and exempts the one that said yes. The safeguards worked. The policy rewards the category that lacks them.


Sources
  1. 1. Chinese Researcher Uses DeepSeek AI to Automate Cyber-Attacks
  2. 2. Trump Administration Finalizes AI Cybersecurity Testing Framework
  3. 3. Trump Administration Declines Safety Testing for Open-Weight AI Models
  4. 4. OpenAI Warns Next-Gen AI Models Pose High Cybersecurity Risk
  5. 5. Trump Signs Executive Order for Voluntary AI Security Vetting
  6. 6. Trump Cancels AI Executive Order After Tech Executive Lobbying
  7. 7. Zhipu AI Releases GLM-5.2 Using Huawei Processors

Keep reading in the app

The full perspective, free in the app.

Download on the App StoreComing soonGoogle Play