HumanEval, developed by OpenAI, is a benchmark of 164 Python problems designed to assess an LLM's code generation and functional correctness, with Claude Sonnet 4 achieving a 95.1% success rate in one study.
Open the full topic