While Claude 3.5 Sonnet scores higher in graduate-level reasoning (GPQA) at 59% compared to GPT-4o's 54%, GPT-4o leads in the MATH benchmark with 76.6% against Claude 3.5 Sonnet's 71.1%.
Open the full topic