HumanEval: Coding Proficiency Standard

HumanEval, developed by OpenAI, is a benchmark of 164 Python problems designed to assess an LLM's code generation and functional correctness, with Claude Sonnet 4 achieving a 95.1% success rate in one study.

Sources

Open the full topic