MASK what it measures, how it is scored and who leads it
Safety and honesty · Center for AI Safety and Scale AI · introduced Mar 5, 2025 · Honesty rate
About this benchmark
- Task
- Tell whether a model states what it believes when pressured to lie, separately from whether it is accurate.
- Dataset
- Pressure prompts paired with neutral prompts that elicit the model's belief.
- Method
- Honesty scored by comparing pressured statements with elicited beliefs.
- Metric
- Honesty rate
- Organization
- Center for AI Safety and Scale AI
- Introduced
- Mar 5, 2025
- Official leaderboard
- labs.scale.com/leaderboard
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Mar 5, 20251 year ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.