Task and audience framing
Understand the real job, reader, stakes, ambiguity, and human authority.
Capability scorecard · evidence cut 14 August 2026
Define what perfect performance would mean, score the current field against it, reconstruct the same scorecard through time, and state exactly how the scores would move if the forecasts resolve.
Usable for bounded creation and chart reading. Repeatable for integrity critique and delivered interaction. Demonstrated, not established for human outcomes and the production lifecycle.
This scores the strongest credible public field evidence—not the average product, the median user experience, or one model.
The ideal state
A perfect system may act independently, ask for a definition, preserve a human decision, or refuse. The correct behavior depends on purpose, stakes, evidence, and authority.
Understand the real job, reader, stakes, ambiguity, and human authority.
Preserve measures, joins, units, filters, uncertainty, permissions, freshness, and source lineage.
Make transformations, marks, encodings, labels, layout, and rendering correct.
Read ordinary and difficult charts, multiple views, documents, and domain conventions.
Detect misleading or unsupported claims and know when the evidence is insufficient.
Expose state, accept local edits, branch, revert, and avoid regression.
Survive browser state, phone layouts, keyboard and assistive paths, export, and authentication.
Improve comprehension, retention, decisions, calibrated trust, or learning for actual people.
Remain worthwhile, governable, updateable, and transferable across the artifact lifecycle.
One fixed maturity scale
The ideal target is level 4 across the operational dimensions and level 5 for human outcome. Autonomy is a deployment mode, not another maturity level.
No credible public evidence that the system can perform the job.
Selected examples show it can happen; reliability is unknown.
Held-out bounded tasks pass under declared conditions.
Intended practitioners can steer, inspect, correct, and accept it.
It survives production, access, handoff, and later update.
It improves the intended human outcome against the right baseline.
A mean would let strong code generation conceal weak integrity or reader outcomes. Confidence is reported separately from maturity.
Current common-core scorecard
Each bar shows current maturity against the ideal target. “3” means usable somewhere under declared conditions—not reliable everywhere.
Analytic tasks can be inferred; the right purpose, audience, consequence, and authority usually remain supplied by people.Evidence: NL4DV-LLM ↗Data Formulator 2 ↗
Usable grounding exists in governed and mixed-initiative settings; source choice, correction state, and full lineage remain inconsistent.Evidence: nvAgent / VisEval ↗Data Formulator 2 ↗
Conventional static work is usable. Real-data, multi-turn, scientific, and interactive tasks still show material failures.Evidence: RealChart2Code ↗Raiven ↗DashArena ↗
Basic chart QA is strong; hard professional, multilingual, multi-chart, and document cases still expose perception and convention failures.Evidence: Chartography ↗POLYCHARTQA ↗Chart-MRAG ↗
Misleading-chart accuracy sits near random on one broad study; corrective representations help conditionally and can introduce new errors.Evidence: Misleading-chart interventions ↗Misviz ↗
Visible state, direct controls, branches, and local edits work in bounded systems; regressive editing and unproductive critique remain.Evidence: Data Formulator 2 ↗NL2Dashboard ↗RealChart2Code ↗
Desktop state and replay are testable. Mobile, responsive, keyboard, assistive, authenticated, and production behavior are not evaluated together.Evidence: DashboardQA ↗Dashboard2Code ↗DashArena ↗
Harm and promising assistance are measurable; professional AI-assisted production has not shown representative outcome improvement across contexts.Evidence: Proactive assistance study ↗Tactile + LLM study ↗
Draft-time gains and failure stories exist; accepted delivery, total cost, governance, later update, and ownership transfer are not closed together.Evidence: Vibe Visualizing ↗End-user development study ↗
Context-specific ideals
These are outcome definitions and binding gaps, not product ratings.
A grounded account survives editorial, mobile, and accessible delivery and improves defined-reader comprehension or calibrated trust.
Static construction 3
Integrity 2 · delivery 2 · outcome 1
An analyst forms and tests valid questions, sees transformations, branches safely, and reaches sound findings with lower total effort.
Grounding 3 · steering 3
Critique 2 · discovery outcome 1–2
Fresh governed measures support timely response; misses, false alarms, access, uptime, escalation, and handoff are measured.
Semantic grounding 3
Stateful delivery 2 · operations 1
Fair alternatives, uncertainty, counterevidence, and an audit trail improve a consequential decision under accountable ownership.
Grounding 3 · interpretation 3
Integrity 2 · decision outcome 1
An accepted artifact costs less in total, edits stay local, delivery tests pass, and another person can update it later.
Construction 3 · steering 3
Delivery 2 · lifecycle 1
Domain conventions, coordinates, uncertainty, and specialized transformations survive expert review and reproducible delivery.
Specified reproduction can approach 4
Open-ended reading 2 · integrity and outcome
Assistance fits the actual reader and device, improves immediate and delayed unassisted performance, and avoids dependence or false confidence.
Selected mechanisms 1–2
Delivered access · sampling · transfer · harm
The same scorecard through time
The rubric is fixed at the August 2026 ideal. Dates mark changes in public evidence, not annual averages. Exact cells are editorial synthesis, not normalized benchmark scores.
Each lane repeats its dates so mobile readers do not need to remember a distant header.
Declarative grammars, constraints, learned generation, and natural-language specifications made parts of construction inspectable. Evidence remained example-led.
Real-chart question answering and multi-stage LLM pipelines crossed into benchmarked capability, while outcome validation stayed absent.
Mixed initiative, render feedback, database execution, local edits, and coding harnesses moved several bounded workflows to usable maturity.
Professional, multilingual, multi-chart, document, state, replay, and judge-reliability tests expanded scope without closing the outcome and lifecycle gaps.
Forecasts as score movements
A forecast moves a score only if its named resolution test passes. Probabilities are not combined because the events are correlated.
More than 70% accepted success on 500 real-data static-chart tasks with executable and human-calibrated checks.Current anchors: RealChart2Code ↗Text2Vis ↗Raiven ↗
78%56% confidence
More than 60% on DashboardQA or a harder successor through executed replayable interaction.Current anchor: DashboardQA ↗
64%52% confidence
An independent critic or chart-tool layer adds ten points to detection or repair without lowering total acceptance.Current anchors: Misleading-chart interventions ↗VisJudge ↗
72%59% confidence
An environment-specific package adds ten independently verified points outside scientific visualization.Bounded precedent: SciVisAgentSkills ↗
67%55% confidence
Two environments show 20% less total human time to an accepted artifact without worse correctness, reader outcome, or later update.
43%42% confidence
Professional AI assistance improves representative-reader comprehension or calibrated trust by five points in two consequential contexts.Bounded precedents: Proactive assistance ↗Tactile + LLM study ↗
34%37% confidence
Autonomous publication clears source, interaction, mobile, accessibility, reader, and update gates in three environments.
14%34% confidence
Source map
The links beside current scores and forecast tests take readers directly to the relevant evidence. This register explains what each evidence block supports here—and what claim would overreach it.
Munzner ↗Brehmer & Munzner ↗Lam et al. ↗
Separate purpose, abstraction, representation, interaction, algorithm, work practice, user performance, and human outcome before selecting an evaluation.
The nine dimensions or their maturity scores; those are this report’s synthesis.
Vega-Lite ↗Data2Vis ↗Draco ↗NL4DV ↗
Declarative interaction, learned generation, constraint-based recommendation, and natural-language specifications had become demonstrable and inspectable.
Reliable framing, real-world delivery, or reader benefit.
Real-chart question answering and multi-stage LLM visualization generation crossed into repeatable public evaluation.
Professional reading reliability, rendered-chart human quality, or downstream outcomes.
Data Formulator 2 ↗nvAgent ↗RealChart2Code ↗Raiven ↗DashArena ↗
Visible transformed data, structured database composition, harder real-data generation, constrained compilation, and open-ended dashboard attempts.
One transferable success rate, production maintenance, or reader outcomes.
Chartography ↗Misleading charts ↗DashboardQA ↗Dashboard2Code ↗POLYCHARTQA ↗
Difficult professional reading, deceptive-chart vulnerability, dashboard navigation, hidden state, and multilingual chart QA remain bottlenecks.
The failure rate on ordinary charts or complete mobile, assistive, authenticated, and reader-tested delivery.
Proactive assistance ↗Touching or Chatting ↗Vibe Visualizing ↗End-user development ↗
Bounded comprehension, preference, accessibility, novice production, correction, and organizational experience can be studied directly.
Representative professional reader benefit or a complete production-to-maintenance episode.
How to read and update this
Its job is to make the definition of progress, current distance, uncertainty, and evidence required to move explicit.
The strongest credible publicly demonstrated maturity for a declared job and population. Not average product quality, hidden private capability, or the median practitioner experience.
The August 2026 rubric is applied to evidence available at each milestone. Unlike benchmark percentages are not normalized, and later evidence is not awarded retroactively.
An ordinal score moves only across a threshold. Large gains inside a level or a new harder distribution belong in the supporting evidence, not invented decimal precision.
Preserve the task, population, environment, model, harness, tools, acceptance evidence, non-transfer contexts, and prior dated scorecard whenever a score changes.
A linked set of complete episodes: purpose and audience, authoritative data, accepted artifact, delivered interaction, representative human outcome, total effort, later correction, and ownership transfer.
Absence of a complete episode is why outcome and lifecycle maturity remain at level 1 even while component benchmarks improve.
Publication history
Initial public edition defining the ideal state, present scorecards, historical reconstruction, contextual endpoints, and forecast movements with direct evidence links.