Skip to content

HumanEval what it measures, how it is scored and who leads it

Coding · OpenAI · introduced Jul 7, 2021 · pass@1 · saturated

About this benchmark

Task
Write a Python function from its signature and docstring.
Dataset
164 hand-written programming problems with unit tests.
Method
Generated code is run against hidden tests; pass@k counts a problem solved if any of k samples passes.
Metric
pass@1
Organization
OpenAI
Introduced
Jul 7, 2021

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the Codex paper.Jul 7, 20215 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.