Terminal-Bench what it measures, how it is scored and who leads it
Coding · Stanford University and Laude Institute · introduced May 19, 2025 · Resolution rate
About this benchmark
- Task
- Complete hard, realistic tasks in a computer terminal: build, debug, configure, analyse.
- Dataset
- Hand-built tasks in containerised terminal environments, each with a test script.
- Method
- An agent works in the container until done or timed out; the task passes if its tests pass.
- Metric
- Resolution rate
- Organization
- Stanford University and Laude Institute
- Introduced
- May 19, 2025
- Official leaderboard
- tbench.ai/
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 4.0Recalibrated task resources, fixed tasks, removed saturated ones. | Aug 28, 20264 weeks ago | Recalibrated task resources, fixed tasks, removed saturated ones. |
| 3.0New frontier tasks. | Jul 30, 20268 weeks ago | New frontier tasks. |
| 2.1Fixes 28 tasks of 2.0. | May 6, 20264 months ago | Fixes 28 tasks of 2.0. |
| 2.089 harder tasks, with the Harbor evaluation package. | Nov 7, 202510 months ago | 89 harder tasks, with the Harbor evaluation package. |
| 1First release with the Terminus agent. | May 19, 20251 year ago | First release with the Terminus agent. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.