July 24, 2026 ← EurekaRaven AI
EurekaRaven AI

Society

How AI guardrails are impeding the work of offensive cybersecurity researchers

6:00 PM · July 24, 2026

Cybersecurity researchers who search for software vulnerabilities and build exploits to fix them before criminals do told TechCrunch that AI safety guardrails, intended to stop malicious hackers, are increasingly getting in their way too. The scrutiny intensified after the US government imposed export control restrictions in June on Anthropic’s Mythos and Fable models over concerns their guardrails could be bypassed for cyberattacks, restrictions that have since been lifted for Fable but only partially for Mythos. Security researcher Mark Dowd said he isn’t comfortable with large AI companies making arbitrary calls about what counts as safe in security, while NCC Group’s chief scientist Chris Anley described the tools as simultaneously essential for defenders and unavoidably usable as weapons, making the guardrails impossible to cleanly separate from legitimate work. Chris Thompson, head of RemoteThreat and founder of Offensive AI Con, said the inconsistency of guardrails from day to day forces researchers to spend more time negotiating with models than actually analyzing vulnerabilities, a dynamic he said is pushing responsible researchers toward unrestricted Chinese open-source models like GLM instead of US-governed systems. Not everyone agrees the guardrails are the core problem: one researcher said he simply doesn’t use AI for the offensive parts of his work at all, relying on it only for reverse engineering and tool-building. The debate adds a new dimension to the broader fight over how tightly frontier AI labs should restrict their most capable models, suggesting that overly blunt guardrails could end up steering serious security researchers toward less accountable alternatives rather than making anyone safer.

Read the full story at techcrunch.com →