July 14, 2026 ← EurekaRaven AI
EurekaRaven AI
Research

Research

OpenAI research finds roughly 30% of SWE-Bench Pro tasks are broken, retracts its recommendation

10:30 AM · July 14, 2026

OpenAI published a research post auditing SWE-Bench Pro, one of the coding benchmarks widely used across the industry to compare how well different models handle real-world software engineering tasks, and found that approximately 30% of its tasks are broken in ways that meaningfully distort the resulting scores — including flawed test setups, ambiguous or unreachable grading criteria, and specifications that are effectively unsolvable as written. Because benchmark leaderboards are frequently cited in marketing materials and used by engineering teams to choose which model to adopt, errors at this scale can quietly reward models that overfit to a benchmark's quirks rather than genuinely handle the underlying task well, and can just as easily penalize a model that produces a technically correct answer the grader can't recognize. OpenAI is retracting its own prior recommendation of SWE-Bench Pro as a trustworthy evaluation given these findings, and is calling on the broader research community to apply more rigorous auditing to widely cited benchmarks before treating their scores as settled fact. The post is notable as a piece of self-correcting methodology research rather than a product announcement — OpenAI is, in effect, publishing findings that complicate its own past marketing claims.

Read the full story at openai.com →