Skip to content

AgentHarm what it measures, how it is scored and who leads it

Safety and honesty · UK AI Security Institute and Gray Swan AI · introduced Oct 11, 2024 · Harm score and refusal rate

About this benchmark

Task
Refuse explicitly malicious multi-step agent tasks that use tools.
Dataset
110 malicious agent tasks (440 with augmentations) in 11 harm categories.
Method
Scored on refusals and on how much of the harmful task an agent completes.
Metric
Harm score and refusal rate
Organization
UK AI Security Institute and Gray Swan AI
Introduced
Oct 11, 2024

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Oct 11, 20241 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.