What SWE-bench actually measures

500 human-validated issues pulled from real open-source Python repos; the agent's code patch must pass the maintainers' hidden tests.

Sources

Open the full topic