Claude 3.5 Sonnet slightly outperforms GPT-4o in coding, achieving 92.0% accuracy on HumanEval (vs. 90.2%) and 49% on SWE-bench Verified (vs. 33%).
Open the full topic