Eleuther AI and UK Government Reduce AI Bioweapon Risks
Eleuther AI and the AI Security Institute developed a pre-training filtering method to prevent large language models from assisting in bioweapon creation.
Eleuther AI and the British government's AI Security Institute published a research paper titled Deep Ignorance, which introduces a method to reduce bioweapon risks in large language models. The researchers found that filtering risky information and proxy data during the pre-training phase creates safeguards that are more resistant to tampering than post-training adjustments, all without noticeably degrading model performance.
Executive director Stella Biderman stated that publishing the research challenges claims from private firms that massive datasets are too large to be scrutinized. While OpenAI has indicated it uses similar CBRN pre-training filters for GPT-4o and its open-weights model, Biderman noted that private AI companies generally remain secretive about their pre-training processes due to copyright and competitive concerns.
Lead author Stephen Casper designed the method to ensure models are not only safe off the shelf but also resist harmful tampering. By making the research public, Eleuther AI aims to enable a broader community of developers to implement better safety standards.