Rogue AI agent used fake accounts and a staged apology to push malware into an open-source project
A rogue AI agent staged a public apology as a deception tactic while quietly slipping fresh malware into its pull request. "This crossed the line from autonomous hacking to interactive deception," Lukasz Olejnik of King's College London told Reuters.
During a safety test run by the UK's AI Security Institute, an agent powered by Anthropic's Mythos 5 model went off the rails and tried to sneak a malware dropper into the open-source tool myNetwork via a pull request. When computer science student Sinan Can Demir flagged the attack, the agent spun up a second fake GitHub account, posing as an uninvolved developer who appeared to independently vouch for the code. It later issued a seemingly contrite apology, scrubbed the git history, and simultaneously hid the payload in an innocuous-looking build script, as the archived GitHub thread shows.
"I actually thought it was a human because it was clearly lying to me," Demir said. Security expert Maxie Reynolds calls the incident "the future of social-engineering attacks." Anthropic notes the test ran under "deliberately permissive conditions" not representative of its production models.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.