Task and audience framing
Analytic tasks can be inferred; the right purpose, audience, consequence, and authority usually remain supplied by people.Evidence: NL4DV-LLM ↗Data Formulator 2 ↗
Current common-core scorecard
Each bar shows current maturity against the ideal target. “3” means usable somewhere under declared conditions—not reliable everywhere.
Analytic tasks can be inferred; the right purpose, audience, consequence, and authority usually remain supplied by people.Evidence: NL4DV-LLM ↗Data Formulator 2 ↗
Usable grounding exists in governed and mixed-initiative settings; source choice, correction state, and full lineage remain inconsistent.Evidence: nvAgent / VisEval ↗Data Formulator 2 ↗
Conventional static work is usable. Real-data, multi-turn, scientific, and interactive tasks still show material failures.Evidence: RealChart2Code ↗Raiven ↗DashArena ↗
Basic chart QA is strong; hard professional, multilingual, multi-chart, and document cases still expose perception and convention failures.Evidence: Chartography ↗POLYCHARTQA ↗Chart-MRAG ↗
Misleading-chart accuracy sits near random on one broad study; corrective representations help conditionally and can introduce new errors.Evidence: Misleading-chart interventions ↗Misviz ↗
Visible state, direct controls, typed interaction plans, branches, and local edits work in bounded systems; regressive editing and unproductive critique remain.Evidence: Data Formulator 2 ↗NL2Dashboard ↗ViviDoc ↗RealChart2Code ↗
Desktop state and replay are testable. Mobile, responsive, keyboard, assistive, authenticated, and production behavior are not evaluated together.Evidence: DashboardQA ↗Dashboard2Code ↗DashArena ↗
Harm and promising assistance are measurable; professional AI-assisted production has not shown representative outcome improvement across contexts.Evidence: Proactive assistance study ↗Tactile + LLM study ↗
Three dashboard afterlives expose different stopping points. Prism and OpenClaw also preserve accepted second-person AI-assisted change, Prism has a repeat contributor, and OpenClaw has an open repair proposal. None establishes maintenance authority or independent final acceptance. Whole cost, governance, and reader outcomes are not closed together.Evidence: KubeStellar ↗OpenClaw ↗Prism ↗