10 reference pages

Benchmark library

Task-and-grader profiles for the evaluations carrying the main capability claims in this research.

Generation and reconstruction · 2024

Plot2Code

A benchmark for reconstructing a scientific plot as executable plotting code from its image.

Open

Generation and reconstruction · 2025

Text2Vis

An end-to-end benchmark combining data, a question, an answer, chart code, and annotated visual evidence.

Open

Generation and reconstruction · 2026

RealChart2Code

A chart-to-code benchmark built from real source data, multi-panel tasks, and iterative refinement.

Open

Chart reading and critique · 2025

ChartQAPro

A realistic chart question-answering benchmark with conversational, hypothetical, fact-checking, and unanswerable cases.

Open

Professional chart reading · 2026

Chartography

A deliberately difficult set of practitioner-authored professional chart-reading tasks.

Open

Multimodal evaluation · 2026

MM-JudgeBench

A multilingual pairwise-preference benchmark for multimodal judges that includes a chart-centric subset.

Open

Interactive dashboards · 2026

Dashboard2Code

A benchmark for reconstructing interactive Plotly Dash applications from screenshots, optional DOM, and interaction.

Open

Interactive dashboard generation · 2026

DashArena

A benchmark, browser executor, and human-calibrated judge for open-ended interactive dashboard generation.

Open