The Model Crown: Benchmarks
The #1 model rotates monthly — Google, OpenAI and Anthropic keep swapping the LMArena and SWE-bench lead.
The top 3
- LMArena: The People's Vote: Human-preference rankings on LMArena flip between Google's Gemini, OpenAI's GPT and xAI's Grok frontier models.
- SWE-bench: The Coding Test: On real GitHub-bug benchmarks, Anthropic's Claude models have repeatedly set the pace for agentic coding.
- The Race Is Nearly Tied: Stanford's AI Index shows top labs separated by low single-digit points — the 'best model' gap keeps shrinking.
Sources
Open the full topic