HumanEval: Code Generation Prowess

HumanEval is a benchmark consisting of 164 Python programming problems designed by OpenAI to assess the functional correctness of LLMs in generating accurate and working code.

Sources

Open the full topic