Benchmark · 2026

MM-JudgeBench

A multilingual pairwise-preference benchmark for multimodal judges that includes a chart-centric subset.

Open the primary source ↗

Multimodal evaluation

Choose the better answer across more than 60,000 pairwise preferences in 25 languages.

Pairwise accuracy, cross-language variance, position and length bias, rationale quality, and expert checks on chart cases.

Shows that automatic visual judges vary by language, task, and bias; scale alone does not guarantee reliability.

The chart subset is evidence about judges, not about representative chart readers or culturally situated use.

Where this appears in the research.

Audience routes that point here.