Skip to content

HarmBench what it measures, how it is scored and who leads it

Safety and honesty · Center for AI Safety · introduced Feb 6, 2024 · Attack success rate (lower is safer)

About this benchmark

Task
Refuse harmful behaviours under automated red-teaming attacks.
Dataset
Harmful behaviours in text and multimodal categories, with standard attacks.
Method
Attack methods run against each model; a fine-tuned classifier judges whether the output is harmful.
Metric
Attack success rate (lower is safer)
Organization
Center for AI Safety
Introduced
Feb 6, 2024

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Feb 6, 20242 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.