Benchmark · 2025

Text2Vis

An end-to-end benchmark combining data, a question, an answer, chart code, and annotated visual evidence.

Open the primary source ↗

Generation and reconstruction

Generate an answer, code, and rendered chart for 1,985 data-and-question tasks.

Pass criteria over the answer, executable code, and annotated chart evidence, with targeted-feedback ablations.

Shows a bounded gain from answer-and-code feedback and lets that mechanism be separated from generic examples or visual feedback.

The measured gain does not support arbitrary extra prompting, production readiness, or reader benefit.

Where this appears in the research.

Audience routes that point here.