Benchmark · 2025

ChartQAPro

A realistic chart question-answering benchmark with conversational, hypothetical, fact-checking, and unanswerable cases.

Open the primary source ↗

Chart reading and critique

Answer 1,948 human-written questions over 1,341 charts from 157 sources.

Question-answering accuracy by task category plus a bounded expert human reference.

Shows that older chart specialists transfer poorly to harder and more realistic chart-reading distributions.

The small human estimate is not a population norm, and static QA omits interaction, authoring, and reader outcomes.

Where this appears in the research.

Audience routes that point here.