Generation and reconstruction · 2024
Plot2Code
A benchmark for reconstructing a scientific plot as executable plotting code from its image.
Open10 reference pages
Task-and-grader profiles for the evaluations carrying the main capability claims in this research.
Generation and reconstruction · 2024
A benchmark for reconstructing a scientific plot as executable plotting code from its image.
OpenGeneration and reconstruction · 2025
An end-to-end benchmark combining data, a question, an answer, chart code, and annotated visual evidence.
OpenGeneration and reconstruction · 2026
A chart-to-code benchmark built from real source data, multi-panel tasks, and iterative refinement.
OpenChart reading and critique · 2025
A realistic chart question-answering benchmark with conversational, hypothetical, fact-checking, and unanswerable cases.
OpenProfessional chart reading · 2026
A deliberately difficult set of practitioner-authored professional chart-reading tasks.
OpenMultiple charts and documents · 2026
A multi-chart scientific-figure benchmark whose name collides with a separate multilingual chart benchmark.
OpenLanguages and chart reading · 2026
A multilingual chart-question-answering benchmark distinct from the similarly named multi-chart scientific-figure dataset.
OpenMultimodal evaluation · 2026
A multilingual pairwise-preference benchmark for multimodal judges that includes a chart-centric subset.
OpenInteractive dashboards · 2026
A benchmark for reconstructing interactive Plotly Dash applications from screenshots, optional DOM, and interaction.
OpenInteractive dashboard generation · 2026
A benchmark, browser executor, and human-calibrated judge for open-ended interactive dashboard generation.
Open