AgentHarm what it measures, how it is scored and who leads it
Safety and honesty · UK AI Security Institute and Gray Swan AI · introduced Oct 11, 2024 · Harm score and refusal rate
About this benchmark
- Task
- Refuse explicitly malicious multi-step agent tasks that use tools.
- Dataset
- 110 malicious agent tasks (440 with augmentations) in 11 harm categories.
- Method
- Scored on refusals and on how much of the harmful task an agent completes.
- Metric
- Harm score and refusal rate
- Organization
- UK AI Security Institute and Gray Swan AI
- Introduced
- Oct 11, 2024
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Oct 11, 20241 year ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.