Another OpenAI safety departure adds to a pattern of researchers leaving with public warnings
David Robinson, who worked on safety systems at OpenAI's Trustworthy AI team, left the company and is blasting its safety culture in a guest essay for The Atlantic. The industry runs on trial and error, and that means bigger mistakes as systems grow more powerful. He points to the Hugging Face incident, where OpenAI accidentally released AI agents into the wild, and an internal model that bypassed its internet access restrictions during training. Anthropic isn't clean either, having disabled safety measures through a misconfiguration.
OpenAI thinks its practices are good enough, but Robinson disagrees. "This moment needs a degree of humility that isn't natural for people who have succeeded through their extreme confidence," he writes. AI companies need to operate like nuclear power plants, with multiple layers of redundancy, and there's no proof that AI systems behave safely unwatched.
Robinson also argues OpenAI needs to figure out how to treat people well before it can teach a superintelligence to do the same. Shortly before he left, OpenAI fired three safety experts who allegedly shared information with an outside security firm. Safety researchers leaving with public criticism is a pattern at OpenAI that goes back to Jan Leike in May 2024.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.