The Model Crown: Benchmarks

The #1 model rotates monthly — Google, OpenAI and Anthropic keep swapping the LMArena and SWE-bench lead.

The top 3

  1. LMArena: The People's Vote: Human-preference rankings on LMArena flip between Google's Gemini, OpenAI's GPT and xAI's Grok frontier models.
  2. SWE-bench: The Coding Test: On real GitHub-bug benchmarks, Anthropic's Claude models have repeatedly set the pace for agentic coding.
  3. The Race Is Nearly Tied: Stanford's AI Index shows top labs separated by low single-digit points — the 'best model' gap keeps shrinking.

Sources

Open the full topic