The Scoreboard: How We'd Even Know

Three yardsticks the debate actually fights over — and the latest scores.

The top 3

  1. ARC-AGI: 87.5%, Then a Faceplant: o3 aced ARC-AGI-1 at 87.5%; the harder ARC-AGI-2 sends frontier models back to single digits.
  2. Turing Test: 73% 'Human': In a 2025 UC San Diego study, judges called GPT-4.5 the human more often than the actual person.
  3. The Crowd's Median Date: Metaculus forecasters have pulled their median general-AI date to years — not decades — away.

Sources

Open the full topic