Skip to content

MMBench what it measures, how it is scored and who leads it

Multimodal · Shanghai AI Laboratory and collaborators · introduced Jul 12, 2023 · Accuracy

About this benchmark

Task
Answer multiple-choice questions testing 20 vision-language abilities.
Dataset
Bilingual (English and Chinese) multiple-choice questions over images.
Method
CircularEval: each question is asked with its options rotated and must be right every time.
Metric
Accuracy
Organization
Shanghai AI Laboratory and collaborators
Introduced
Jul 12, 2023

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Jul 12, 20233 years ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.