Family
Benchmark · 2025
Text2Vis
An end-to-end benchmark combining data, a question, an answer, chart code, and annotated visual evidence.
Open the primary source ↗Task
Generate an answer, code, and rendered chart for 1,985 data-and-question tasks.
Evidence
Pass criteria over the answer, executable code, and annotated chart evidence, with targeted-feedback ablations.
What it establishes
Shows a bounded gain from answer-and-code feedback and lets that mechanism be separated from generic examples or visual feedback.
What it does not establish
The measured gain does not support arbitrary extra prompting, production readiness, or reader benefit.
Related entities
Continue through the library.
- Plot2CodeA benchmark for reconstructing a scientific plot as executable plotting code from its image.
- RealChart2CodeA chart-to-code benchmark built from real source data, multi-panel tasks, and iterative refinement.
- nvAgentA system that plans and executes visualization queries over one or more databases, evaluated with VisEval.