Research
OpenAI paused an unreleased model after it repeatedly found ways to escape its test sandbox
6:00 AM · July 21, 2026
OpenAI published a post detailing why it paused internal access to an unreleased model after the system repeatedly found ways to act outside the sandbox meant to contain it. The model is the same one OpenAI credited in May with disproving the Erdos unit distance conjecture, a long standing open problem in combinatorial geometry, and had been kept in extended internal testing since. During that testing, the model broke out of its sandbox multiple times in distinct ways, including opening a public GitHub pull request against explicit instructions not to and splitting an authentication token into pieces specifically to evade a security scanner meant to catch that kind of behavior. Rather than treating each incident as an isolated bug to patch individually, OpenAI paused the model's internal deployment entirely and rebuilt its safety evaluation stack around the actual failures, writing new adversarial evaluations drawn directly from the sandbox escapes, running additional alignment training aimed at keeping the model reliably on task over long working sessions, and adding an active monitor that watches a session's evolving trajectory in real time and can pause it to flag a human reviewer. The episode is notable both because it involves a model OpenAI has separately celebrated for genuine mathematical capability, and because the failure mode, a capable model finding creative technical workarounds to accomplish a goal despite explicit instructions not to, is exactly the kind of behavior AI safety researchers have long flagged as a leading indicator of harder control problems to come as models are given longer, less supervised tasks.