OpenAI's GPT-5.6 Sol and Anthropic's Claude Mythos Preview both lead the GPQA benchmark with 94.6% as of August 2026, showcasing their advanced reasoning in specialized scientific domains.
Open the full topic