Family
Benchmark · 2026
MM-JudgeBench
A multilingual pairwise-preference benchmark for multimodal judges that includes a chart-centric subset.
Open the primary source ↗Task
Choose the better answer across more than 60,000 pairwise preferences in 25 languages.
Evidence
Pairwise accuracy, cross-language variance, position and length bias, rationale quality, and expert checks on chart cases.
What it establishes
Shows that automatic visual judges vary by language, task, and bias; scale alone does not guarantee reliability.
What it does not establish
The chart subset is evidence about judges, not about representative chart readers or culturally situated use.
Related entities
Continue through the library.
- POLYCHARTQA: multilingual chart QAA multilingual chart-question-answering benchmark distinct from the similarly named multi-chart scientific-figure dataset.
- PolyChartQA: multi-chart figuresA multi-chart scientific-figure benchmark whose name collides with a separate multilingual chart benchmark.
- ChartographyA deliberately difficult set of practitioner-authored professional chart-reading tasks.