SWE-bench: The Coding Test

On real GitHub-bug benchmarks, Anthropic's Claude models have repeatedly set the pace for agentic coding.

Sources

Open the full topic