HumanEval what it measures, how it is scored and who leads it
Coding · OpenAI · introduced Jul 7, 2021 · pass@1 · saturated
About this benchmark
- Task
- Write a Python function from its signature and docstring.
- Dataset
- 164 hand-written programming problems with unit tests.
- Method
- Generated code is run against hidden tests; pass@k counts a problem solved if any of k samples passes.
- Metric
- pass@1
- Organization
- OpenAI
- Introduced
- Jul 7, 2021
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the Codex paper. | Jul 7, 20215 years ago | Released with the Codex paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.