Benchmark Performance Leaders

GPT-5.5 achieved an 82.7% score on Terminal-Bench 2.0, surpassing Claude Opus 4.7, while Gemini 3 Pro (with Deep Think) leads on GPQA Diamond at 93.8% and ARC-AGI-2 with 45.1% (with code execution), outperforming GPT-5.1 on both.

Sources

Open the full topic