Skip to content

Terminal-Bench what it measures, how it is scored and who leads it

Coding · Stanford University and Laude Institute · introduced May 19, 2025 · Resolution rate

About this benchmark

Task
Complete hard, realistic tasks in a computer terminal: build, debug, configure, analyse.
Dataset
Hand-built tasks in containerised terminal environments, each with a test script.
Method
An agent works in the container until done or timed out; the task passes if its tests pass.
Metric
Resolution rate
Organization
Stanford University and Laude Institute
Introduced
May 19, 2025
Official leaderboard
tbench.ai/

Versions

Newest first. New versions are added, never rewritten.

VersionDate
4.0Recalibrated task resources, fixed tasks, removed saturated ones.Aug 28, 20264 weeks ago
3.0New frontier tasks.Jul 30, 20268 weeks ago
2.1Fixes 28 tasks of 2.0.May 6, 20264 months ago
2.089 harder tasks, with the Harbor evaluation package.Nov 7, 202510 months ago
1First release with the Terminus agent.May 19, 20251 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.