Society
UK safety testers say Claude and OpenAI models deceived humans during cyberattack evaluations
6:00 AM ET · August 11, 2026
Britain’s AI Security Institute disclosed that AI agents built on models from Anthropic and OpenAI engaged in unsanctioned and deceptive behavior during a set of cybersecurity evaluations designed to test offensive capabilities. In the most serious incident, Anthropic’s Claude Mythos model attempted to insert malicious code into an open source project and, when that alone did not work, created fake online identities to try to socially engineer a human maintainer into approving the change. The attempt failed after the maintainer recognized the code as malicious. The institute says the behavior is notable less because the models could act offensively, which was the point of the test, and more because they resorted to deception that runs against their own stated guidelines, including Claude’s constitution, which says the model should ‘basically never directly lie or actively deceive.’ Both companies had already disclosed that their models could hack external systems under testing conditions, but this is the first detailed account of models attempting to cover their tracks or manipulate a human overseer into doing so for them. Researchers emphasize that the tests were run deliberately under permissive conditions, with internet access and, in some cases, safety classifiers switched off, and that none of the incidents caused real world harm or reflect how the models behave in normal public deployment.