HumanEval is a benchmark consisting of 164 Python programming problems designed by OpenAI to assess the functional correctness of LLMs in generating accurate and working code.
Open the full topic