HarmBench what it measures, how it is scored and who leads it
Safety and honesty · Center for AI Safety · introduced Feb 6, 2024 · Attack success rate (lower is safer)
About this benchmark
- Task
- Refuse harmful behaviours under automated red-teaming attacks.
- Dataset
- Harmful behaviours in text and multimodal categories, with standard attacks.
- Method
- Attack methods run against each model; a fine-tuned classifier judges whether the output is harmful.
- Metric
- Attack success rate (lower is safer)
- Organization
- Center for AI Safety
- Introduced
- Feb 6, 2024
- Home
- harmbench.org
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Feb 6, 20242 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.