The Benchmark Crown: SWE-bench Verified

One number fuels the arguments: the % of 500 real GitHub issues an agent can fix entirely on its own.

The top 3

  1. From low double-digits to 70%+ in ~18 months: Verified scores leapt from roughly the 20% range in 2024 to over 70% by 2025 as agent scaffolds and frontier models matured.
  2. What SWE-bench actually measures: 500 human-validated issues pulled from real open-source Python repos; the agent's code patch must pass the maintainers' hidden tests.
  3. The contamination and gaming debate: Critics flag train/test leakage and scaffolds overfit to the benchmark — why a chart-topping score doesn't equal day-to-day usefulness.

Sources

Open the full topic