The AI Labs Are Handing You the Risk and Keeping the Upside
The safety push — sincere or not — is building a system where the government approves, insures, and runs the most dangerous models, while the labs keep the money and the race.
In February the U.S. government designated Anthropic a national security supply-chain risk, ordered federal agencies to stop using its technology, and called the company woke and radical-left for refusing to strip the guardrails off its models [1]. By June the president was ordering every agency to stop using Anthropic's technology outright [2]. All the while, the NSA had been running Anthropic's most dangerous model — Mythos, the one that found a 27-year-old Linux kernel flaw and simulated a full network takeover — for red-teaming, after the agency itself sought access to it [3][4]. The government was simultaneously the regulator punishing the lab, the customer using its product, and the party whose demand for unrestricted access had created the fight in the first place. That paradox is not a glitch. It is the shape of the whole arrangement now settling into place. The labs keep asking the government to take on more of the risk of AI, and the government keeps saying yes, in three different ways. First, certification. Dario Amodei has proposed an FAA-style framework: frontier models, like airplanes, should pass technical testing and auditing before release, and be blocked or reversed if they fail [5]. The appeal of the airplane analogy is precise, and it is not only about safety. Once a plane passes FAA certification, liability for undiscovered flaws shifts partly toward the regulator that approved it, not just the manufacturer. The shift is real but incomplete — Boeing still faced massive lawsuits after the 737 MAX crashes despite certification [5]. Still, a government-approved process creates a presumption of safety, and a presumption is a useful thing to hold when the lawsuits arrive. Second, insurance. The private market has already made its decision. W.R. Berkley, Chubb, and AIG are asking regulators for permission to exclude AI liabilities from corporate policies entirely, citing the unpredictability of the technology and the threat of a single model failure triggering thousands of simultaneous claims [6]. When the insurers walk away, the only entity left big enough to absorb a correlated loss is the government — the implicit backstop, whether anyone voted for it or not. Third, the fine print. Google, OpenAI, and Anthropic all disclaim responsibility for the accuracy of their models in their terms of service, while 51% of workers believe they would be personally liable for financial harm caused by bad AI output [7]. The lab contracts out of the risk; the person at the desk is left holding it. None of this requires the labs to be cynical. The safety concern is real. Anthropic refused to remove its guardrails for the Pentagon, lost a $200 million contract, and was blacklisted for it [1]. Amodei was blunt about why.
Such use cases have never been included in our contracts with the Department of War, and we believe they should not be included now. — Dario Amodei
That is a genuine sacrifice, not a gesture. But the architecture the sacrifice produces does not depend on anyone's motives. The labs keep the upside — a $2 trillion IPO and $65 billion in annualized revenue [8] — while the government absorbs the risk of approval, the risk of insurance, and the risk of actually running the models. The critics have said this out loud. Dan Hendrycks calls the industry's safety work "performative safe paperwork or controlled opposition designed to maintain development momentum" [9]. Bella Devaan puts it as a structural point.
They can’t be responsible for their own regulation or for their own definition of what their altruistic mandate is. — Bella Devaan
David Sacks, the former AI czar, has the political version.
Dario believes frontier AI is too powerful to distribute; we believe it is too powerful to centralize. — David Sacks
The one thing the safety push has never asked the labs to surrender is the race itself. Anthropic is pushing toward AI writing 80% of its own code; OpenAI targets full automation of its researchers by 2028 [10]. Asked whether the world should slow down, Anthropic co-founder Jack Clark said no.
We actually believe this should be slowed down … We need some sort of international norm to be able to control this. — Jakub Pachocki
So the walls line up like this: the government gets the approval authority, the insurance burden, and the operational risk of running the most dangerous models. The user gets the liability. The labs get the money and the permission to keep building. Nobody designed it that way. It is just where the risk has settled.
- 1. Anthropic Sues Trump Administration Over National Security Blacklist
- 2. US Blocks Anthropic AI Models Over National Security Risks
- 3. OpenAI and Anthropic Launch Specialized AI Cybersecurity Models
- 4. Anthropic Considers $900 Billion Valuation and October IPO
- 5. Anthropic CEO Urges Binding Regulations to Block Dangerous AI Models
- 6. US Insurers Seek to Exclude AI Liabilities From Policies
- 7. Reports Warn of Cognitive Atrophy and AI Verification Deficits
- 8. Anthropic Delays IPO Targeting Record $2 Trillion Valuation
- 9. AI Companies Accelerate Superintelligence Despite Extinction Warnings
- 10. Anthropic and OpenAI Race Toward Recursive AI Self-Improvement