MMBench what it measures, how it is scored and who leads it
Multimodal · Shanghai AI Laboratory and collaborators · introduced Jul 12, 2023 · Accuracy
About this benchmark
- Task
- Answer multiple-choice questions testing 20 vision-language abilities.
- Dataset
- Bilingual (English and Chinese) multiple-choice questions over images.
- Method
- CircularEval: each question is asked with its options rotated and must be right every time.
- Metric
- Accuracy
- Organization
- Shanghai AI Laboratory and collaborators
- Introduced
- Jul 12, 2023
Versions
Newest first. New versions are added, never rewritten.
| Version | Date | |
|---|---|---|
| 1Released with the paper. | Jul 12, 20233 years ago | Released with the paper. |
Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.