The Benchmark Crown: SWE-bench Verified
One number fuels the arguments: the % of 500 real GitHub issues an agent can fix entirely on its own.
The top 3
- From low double-digits to 70%+ in ~18 months: Verified scores leapt from roughly the 20% range in 2024 to over 70% by 2025 as agent scaffolds and frontier models matured.
- What SWE-bench actually measures: 500 human-validated issues pulled from real open-source Python repos; the agent's code patch must pass the maintainers' hidden tests.
- The contamination and gaming debate: Critics flag train/test leakage and scaffolds overfit to the benchmark — why a chart-topping score doesn't equal day-to-day usefulness.
Sources
Open the full topic