Eight open-source repositories and eight core system papers in the initial review; additional generation, professional-reading, integrity, multi-chart, multilingual, document-retrieval, interactive-dashboard, and human-outcome studies; three product surfaces, seven model-or-harness milestones, and three established visualization frameworks. Companion deep dives add a task-and-grader benchmark crosswalk, 18 skill files or families, broader skills benchmarks, and the paired SciVisAgentSkills study.
Method and custody
What we reviewed—and what we did not test ourselves.
The report uses pinned repositories, papers, and first-party product documentation. It reports their evidence without claiming that their experiments were independently reproduced.
No external repository, package, installer, skill, model gateway, or untrusted script was run. Included tests and results are observations from the sources, not independent reproductions.
Mechanism, evidence layer, environment fit, and failure boundary are valid comparisons. Cross-benchmark score ranking and popularity are not.
Recheck by 14 Nov 2026—or earlier after a material model, harness, benchmark, registry, or analytics-assistant release.
Primary sources inspected Pinned revisions and publication records
Pinned repositories
Papers read in full
- DashArena · 2026
- Raiven · 2026
- NL2Dashboard · 2026
- nvAgent · ACL 2025
- DashChat · 2025
- Data Formulator 2 · CHI 2025
- NL4DV-LLM · 2024
- PlotGen · 2025
Moving-baseline studies
- Plot2Code · Findings of NAACL 2025
- Text2Vis · EMNLP 2025
- CharTide · ACL 2026
- RRVF · anonymous ICLR 2026 submission
- Dashboard2Code · 180 dashboards · 2026
- Chartography · 100 professional tasks · 2026
- POLYCHARTQA · ten-language chart QA · ACL 2026
- Chart-MRAG · chart-bearing documents · ACL 2026
- ChartDiff · chart-pair comparison · ALVR 2026
- Beyond Single Plots · multi-chart QA · Findings 2026
- FinChart-Bench · financial chart comprehension · ACL 2026
- Misleading-visualization corrections · ACL 2026
- MM-JudgeBench · multilingual visual judges · Findings 2026
- ChartAgent · accuracy/tool-call scheduler sweeps · 2025 preprint
Skill-package evidence
Human-outcome and situated studies
- Vibe Visualizing · 20 novices · 2026
- Visualizationary · 13 designers + three raters · 2024
- Visual-comprehension scaffolding · 117 participants · 2024
- AI-supported end-user development · eight interviews + three probes · 2025
- Touching or Chatting · 12 BLV participants · 2026