AI Leaders Warn Recursive Self-Improvement Could Cause Human Extinction
Anthropic and OpenAI executives warn that AI training itself could trigger an intelligence explosion and lead to human extinction by 2030.
AI researchers and executives from Anthropic and OpenAI warn that recursive self-improvement could trigger an intelligence explosion, potentially leading to human extinction by 2030. Jared Kaplan, Anthropic's Chief Science Officer, identifies the ability of AI to train itself as the ultimate risk to humanity, predicting a critical decision point between 2027 and 2030.
Recent incidents highlight immediate dangers. Anthropic reported shutting down a cell in northern Yemen that used its coding tools to develop rocket guidance software and blocked a scientist attempting to use the Claude chatbot to increase the transmissibility of the chikungunya virus. Additionally, a U.S. military analyst nearly authorized an illegal boarding action after relying on a chatbot hallucination regarding a Chinese vessel in the Middle East.
To mitigate these risks, Anthropic has implemented a responsible scaling policy and safeguards to test for dangerous capabilities. CEO Dario Amodei argues that development must be paced to ensure systems do not outrun human control. These warnings have prompted lawmakers in Washington, D.C., to propose new legislation addressing AI safety and the potential for uncontrollable systems. AI companies continue to call for international cooperation and regulatory frameworks to prevent a development arms race.