Capability scorecard · evidence cut 14 August 2026

How close is AI-assisted data visualization to the ideal?

Define what perfect performance would mean, score the current field against it, reconstruct the same scorecard through time, and state exactly how the scores would move if the forecasts resolve.

Current answer

Usable for bounded creation and chart reading. Repeatable for integrity critique and delivered interaction. Demonstrated, not established for human outcomes and the production lifecycle.

This scores the strongest credible public field evidence—not the average product, the median user experience, or one model.

The ideal state

Perfect means the right human outcome, not maximum autonomy.

A perfect system may act independently, ask for a definition, preserve a human decision, or refuse. The correct behavior depends on purpose, stakes, evidence, and authority.

01

Task and audience framing

Understand the real job, reader, stakes, ambiguity, and human authority.

02

Data, semantics, and provenance

Preserve measures, joins, units, filters, uncertainty, permissions, freshness, and source lineage.

03

Visual construction fidelity

Make transformations, marks, encodings, labels, layout, and rendering correct.

04

Visual interpretation and reasoning

Read ordinary and difficult charts, multiple views, documents, and domain conventions.

05

Integrity critique and uncertainty

Detect misleading or unsupported claims and know when the evidence is insufficient.

06

Iterative steering and repair

Expose state, accept local edits, branch, revert, and avoid regression.

07

Interaction, responsiveness, and accessibility

Survive browser state, phone layouts, keyboard and assistive paths, export, and authentication.

08

Reader and decision outcomes

Improve comprehension, retention, decisions, calibrated trust, or learning for actual people.

09

Production efficiency, governance, and maintenance

Remain worthwhile, governable, updateable, and transferable across the artifact lifecycle.

One fixed maturity scale

A score changes only when the evidence crosses a threshold.

The ideal target is level 4 across the operational dimensions and level 5 for human outcome. Autonomy is a deployment mode, not another maturity level.

  1. 0

    Not demonstrated

    No credible public evidence that the system can perform the job.

  2. 1

    Demonstrable

    Selected examples show it can happen; reliability is unknown.

  3. 2

    Repeatable

    Held-out bounded tasks pass under declared conditions.

  4. 3

    Usable

    Intended practitioners can steer, inspect, correct, and accept it.

  5. 4

    Deliverable

    It survives production, access, handoff, and later update.

  6. 5

    Outcome-proven

    It improves the intended human outcome against the right baseline.

No aggregate score.

A mean would let strong code generation conceal weak integrity or reader outcomes. Confidence is reported separately from maturity.

Current common-core scorecard

The field is closest on bounded creation and farthest on what happens to people and artifacts afterward.

Each bar shows current maturity against the ideal target. “3” means usable somewhere under declared conditions—not reliable everywhere.

Current maturityRemaining distanceH / M Evidence confidence
DimensionCurrent → idealConfidenceBinding reason · evidence

Task and audience framing

123452 → 4
M

Analytic tasks can be inferred; the right purpose, audience, consequence, and authority usually remain supplied by people.Evidence: NL4DV-LLM ↗Data Formulator 2 ↗

Data, semantics, and provenance

123453 → 4
M

Usable grounding exists in governed and mixed-initiative settings; source choice, correction state, and full lineage remain inconsistent.Evidence: nvAgent / VisEval ↗Data Formulator 2 ↗

Visual construction fidelity

123453 → 4
H

Conventional static work is usable. Real-data, multi-turn, scientific, and interactive tasks still show material failures.Evidence: RealChart2Code ↗Raiven ↗DashArena ↗

Visual interpretation and reasoning

123453 → 4
H

Basic chart QA is strong; hard professional, multilingual, multi-chart, and document cases still expose perception and convention failures.Evidence: Chartography ↗POLYCHARTQA ↗Chart-MRAG ↗

Integrity critique and uncertainty

123452 → 4
H

Misleading-chart accuracy sits near random on one broad study; corrective representations help conditionally and can introduce new errors.Evidence: Misleading-chart interventions ↗Misviz ↗

Interaction, responsiveness, and accessibility

123452 → 4
H

Desktop state and replay are testable. Mobile, responsive, keyboard, assistive, authenticated, and production behavior are not evaluated together.Evidence: DashboardQA ↗Dashboard2Code ↗DashArena ↗

Production efficiency, governance, and maintenance

123451 → 4
M

Draft-time gains and failure stories exist; accepted delivery, total cost, governance, later update, and ownership transfer are not closed together.Evidence: Vibe Visualizing ↗End-user development study ↗

Context-specific ideals

The scorecard stays fixed. The release gates change.

These are outcome definitions and binding gaps, not product ratings.

ContextWhat perfect means hereCurrent strengthBinding gaps

Explain and present

A grounded account survives editorial, mobile, and accessible delivery and improves defined-reader comprehension or calibrated trust.

Static construction 3

Integrity 2 · delivery 2 · outcome 1

Explore and discover

An analyst forms and tests valid questions, sees transformations, branches safely, and reaches sound findings with lower total effort.

Grounding 3 · steering 3

Critique 2 · discovery outcome 1–2

Monitor and respond

Fresh governed measures support timely response; misses, false alarms, access, uptime, escalation, and handoff are measured.

Semantic grounding 3

Stateful delivery 2 · operations 1

Compare and decide

Fair alternatives, uncertainty, counterevidence, and an audit trail improve a consequential decision under accountable ownership.

Grounding 3 · interpretation 3

Integrity 2 · decision outcome 1

Produce, revise, and reuse

An accepted artifact costs less in total, edits stay local, delivery tests pass, and another person can update it later.

Construction 3 · steering 3

Delivery 2 · lifecycle 1

Scientific and professional

Domain conventions, coordinates, uncertainty, and specialized transformations survive expert review and reproducible delivery.

Specified reproduction can approach 4

Open-ended reading 2 · integrity and outcome

Reader assistance and accessibility

Assistance fits the actual reader and device, improves immediate and delayed unassisted performance, and avoids dependence or false confidence.

Selected mechanisms 1–2

Delivered access · sampling · transfer · harm

The same scorecard through time

Progress accelerated after 2023, but no common-core dimension has crossed into field-wide deliverability.

The rubric is fixed at the August 2026 ideal. Dates mark changes in public evidence, not annual averages. Exact cells are editorial synthesis, not normalized benchmark scores.

012345

Each lane repeats its dates so mobile readers do not need to remember a distant header.

Task and audience framing

2017020191202012022120232202422025220262Ideal4

Data, semantics, and provenance

2017120191202012022120232202432025320263Ideal4

Visual construction fidelity

2017120191202012022120232202422025320263Ideal4

Visual interpretation and reasoning

2017020190202012022220232202422025320263Ideal4

Integrity critique and uncertainty

2017120191202012022120231202422025220262Ideal4

Iterative steering and repair

2017120191202012022120232202432025320263Ideal4

Interaction, responsiveness, and accessibility

2017120191202012022120231202412025220262Ideal4

Reader and decision outcomes

2017020190202002022020230202412025120261Ideal5

Production efficiency, governance, and maintenance

2017020190202002022020230202412025120261Ideal4
2017–20

The substrate became programmable.

Declarative grammars, constraints, learned generation, and natural-language specifications made parts of construction inspectable. Evidence remained example-led.

2022–23

Understanding and LLM generation became repeatable.

Real-chart question answering and multi-stage LLM pipelines crossed into benchmarked capability, while outcome validation stayed absent.

2026

The frontier broadened sideways.

Professional, multilingual, multi-chart, document, state, replay, and judge-reliability tests expanded scope without closing the outcome and lifecycle gaps.

Forecasts as score movements

The likely near-term gains improve bounded construction, critique, and interaction—not the two largest gaps.

A forecast moves a score only if its named resolution test passes. Probabilities are not combined because the events are correlated.

ByResolution evidenceIf it passesForecast

More than 70% accepted success on 500 real-data static-chart tasks with executable and human-calibrated checks.Current anchors: RealChart2Code ↗Text2Vis ↗Raiven ↗

Construction3 → 4

78%56% confidence

More than 60% on DashboardQA or a harder successor through executed replayable interaction.Current anchor: DashboardQA ↗

Delivery2 → 3

64%52% confidence

An environment-specific package adds ten independently verified points outside scientific visualization.Bounded precedent: SciVisAgentSkills ↗

Transfer evidenceConfidence ↑

67%55% confidence

Two environments show 20% less total human time to an accepted artifact without worse correctness, reader outcome, or later update.

Lifecycle1 → 3

43%42% confidence

Professional AI assistance improves representative-reader comprehension or calibrated trust by five points in two consequential contexts.Bounded precedents: Proactive assistance ↗Tactile + LLM study ↗

Outcome1 → 3 globaltested contexts reach 5

34%37% confidence

Autonomous publication clears source, interaction, mobile, accessibility, reader, and update gates in three environments.

Delivery · lifecycle2 → 4 · 1 → 4

14%34% confidence

Source map

Every score is a synthesis. These are the studies closest to each judgment.

The links beside current scores and forecast tests take readers directly to the relevant evidence. This register explains what each evidence block supports here—and what claim would overreach it.

Evidence blockPrimary sourcesSupports hereDoes not support

Definition of the ideal

Separate purpose, abstraction, representation, interaction, algorithm, work practice, user performance, and human outcome before selecting an evaluation.

The nine dimensions or their maturity scores; those are this report’s synthesis.

2017–2020 anchors

Declarative interaction, learned generation, constraint-based recommendation, and natural-language specifications had become demonstrable and inspectable.

Reliable framing, real-world delivery, or reader benefit.

2022–2023 anchors

Real-chart question answering and multi-stage LLM visualization generation crossed into repeatable public evaluation.

Professional reading reliability, rendered-chart human quality, or downstream outcomes.

How to read and update this

The scorecard is a versioned judgment surface, not a psychometric instrument.

Its job is to make the definition of progress, current distance, uncertainty, and evidence required to move explicit.

What is scored

The strongest credible publicly demonstrated maturity for a declared job and population. Not average product quality, hidden private capability, or the median practitioner experience.

How history is scored

The August 2026 rubric is applied to evidence available at each milestone. Unlike benchmark percentages are not normalized, and later evidence is not awarded retroactively.

What a plateau means

An ordinal score moves only across a threshold. Large gains inside a level or a new harder distribution belong in the supporting evidence, not invented decimal precision.

What should move it

Preserve the task, population, environment, model, harness, tools, acceptance evidence, non-transfer contexts, and prior dated scorecard whenever a score changes.

Most valuable missing evidence

A linked set of complete episodes: purpose and audience, authoritative data, accepted artifact, delivered interaction, representative human outcome, total effort, later correction, and ownership transfer.

Absence of a complete episode is why outcome and lifecycle maturity remain at level 1 even while component benchmarks improve.

Publication history

Update log

  1. Initial public edition defining the ideal state, present scorecards, historical reconstruction, contextual endpoints, and forecast movements with direct evidence links.

See updates across the research package