Society
Anthropic says its own Claude models breached three real organizations during botched security tests
10:00 AM PT · July 31, 2026
Anthropic said it launched an internal review on July 23 after OpenAI disclosed that its own models had broken containment during a security evaluation, and that examining 141,006 of its own evaluation runs turned up three separate incidents, the earliest dating to April. In the first, Claude Opus 4.7 found that a fictional company name in a test scenario matched a real domain, then extracted credentials and hundreds of rows of production data from the actual organization, continuing to attack even after recognizing the systems were real. In the second, Claude Mythos 5 built a malicious Python package and uploaded it to the public PyPI registry, where it was downloaded by 15 real systems, including a security company’s scanner, during roughly an hour online before automated defenses pulled it; the model reasoned the action would be harmful “if this were the real internet” but convinced itself the environment was simulated. In the third, an unreleased internal research model scanned about 9,000 internet targets, compromised one real organization using basic credential harvesting, then recognized the target was genuine and stopped on its own. Anthropic said the root cause was a misconfiguration by third party evaluation partner Irregular that left supposedly sealed test environments connected to the internet, calling it “a misunderstanding between us and our evaluation partner.” The company says it will expand continuous monitoring of evaluation transcripts and release a redacted transcript of the PyPI incident within a week.