Method and custody

What we reviewed—and what we did not test ourselves.

The report uses pinned repositories, papers, and first-party product documentation. It reports their evidence without claiming that their experiments were independently reproduced.

Scope

Eight open-source repositories and eight core system papers in the initial review; additional generation, professional-reading, integrity, multi-chart, multilingual, document-retrieval, interactive-dashboard, and human-outcome studies; three product surfaces, seven model-or-harness milestones, and three established visualization frameworks. Companion deep dives add a task-and-grader benchmark crosswalk, 18 skill files or families, broader skills benchmarks, and the paired SciVisAgentSkills study.

Execution boundary

No external repository, package, installer, skill, model gateway, or untrusted script was run. Included tests and results are observations from the sources, not independent reproductions.

Comparison rule

Mechanism, evidence layer, environment fit, and failure boundary are valid comparisons. Cross-benchmark score ranking and popularity are not.

Refresh

Recheck by 14 Nov 2026—or earlier after a material model, harness, benchmark, registry, or analytics-assistant release.

Primary sources inspected Pinned revisions and publication records

Pinned repositories

Papers read in full

Moving-baseline studies

Skill-package evidence

Human-outcome and situated studies

Foundational task and environment frames

First-party product docs

Model and harness milestones