# Packaging “The state of AI-assisted data visualization”

Status: editorial architecture, 14 August 2026.

## Recommendation

Treat the work as one evidence base with two renderings:

1. **A comprehensive state-of-the-field report** for readers who need the whole
   argument, shared definitions, comparison rules, and research agenda.
2. **A non-linear series of self-contained essays** for readers who enter
   through one practical question: what models can do, which techniques work,
   what practitioners use, why readers misread, or which gaps matter.

The long report should not be a bundle of the essays pasted together. It should
provide the connective tissue: one scope, one vocabulary, one evidence model,
one timeline, and explicit transitions from capability to system design to
practice to reader outcome. The essays should use the same evidence and figures
but repeat enough local context that no reader needs to remember a definition
from eight chapters earlier.

This supports several audiences without pretending they need the same reading
path:

| Reader | Likely question | Best entry |
| --- | --- | --- |
| Executive, editor, or research commissioner | What is true now, what remains risky, and what deserves investment? | executive summary, convergence, gap ledger, research agenda |
| Visualization or data practitioner | Where does AI remove work, where does it add work, and which surface fits my job? | jobs, tool families, creator experience, evaluation protocol |
| Product or engineering team | Which mechanisms are durable, which scaffolds are decaying, and what should we test? | technique spine, architectures, skills, critics, benchmark crosswalk |
| Researcher | What has actually been measured, where do studies conflict, and what would change the synthesis? | methods, benchmark crosswalk, negative findings, gap ledger, source index |
| Reader, educator, accessibility specialist, or journalist | What happens after the chart is made? | reader outcomes, disclosure and trust, mobile/accessibility, context chapters |

### The current pieces and their roles

- [The state of AI-assisted data visualization research](/reports/research-review/)
  supplies the capability timeline, research synthesis, technique comparison,
  evidence ladder, and research gap ledger.
- [The practitioner and reader experience](/reports/practitioner-and-reader-experience/)
  supplies jobs, tool families, observed adoption, creator experience, reader
  outcomes, production constraints, and its separate gap ledger.
- [Human skills and banked gains](/reports/human-skills-and-gains/)
  supplies the common visualization-literacy framework, a multi-dimensional
  expertise model, evidence by learner and practitioner profile, a strict
  outcome ledger, observed harms, and delayed-learning and field-study designs.
- [Learning data visualization when AI can make the chart](/reports/learning-with-ai/)
  follows the learner from the recent past through August 2026, distinguishes
  productive difficulty from removable friction, and allocates practice across
  judgment, procedural fluency, and skills likely to decay in value.
- [Data-visualization skills for coding agents](/reports/agent-skills/)
  supplies the registry scan, package-role and technique comparisons, dated
  adoption proxies, direct scientific-visualization evidence, and an ablation
  agenda for additional instructions.
- [Specialized vision models for data visualization](/reports/specialized-vision-models/)
  separates six perception and critique roles, compares specialists with
  current general models only on shared tests, and defines where a specialist
  is a useful sensor rather than a final judge.
- [What would unlock the next capabilities in AI-assisted data
  visualization?](/reports/capability-forecast/)
  maps remaining capability gates by environment, separates raw-model and
  technique-led gains, and opens seven dated forecasts with explicit signposts,
  falsifiers, and resolution surfaces.
- [The executive summary](/reports/executive-summary/)
  now supplies chapter 0 across both gap ledgers and the two specialist deep
  dives.
- [The audience, decisions, and rollout strategy](/reports/audience-decisions-and-rollout/)
  maps the primary reader groups, their decision moments and language, the
  public artifact family, permissioned distribution routes, and the evidence
  needed to distinguish visibility from actual use.
- [What AI-assisted data visualization benchmarks actually
  measure](/reports/benchmark-crosswalk/)
  supplies the benchmark measurement ladder, family-by-family crosswalk,
  construction effects, demonstrated metric failures, and the specification
  for a still-missing complete episode.
- [How close is AI-assisted data visualization to the
  ideal?](/reports/ideal-state-and-scorecards/)
  defines a shared ideal spine and context-specific end states, scores nine
  current dimensions without a composite, reconstructs the same ordinal
  scorecard from 2017 through August 2026, and expresses each dated forecast as
  a score movement conditional on its resolution evidence.
- [The content map](/content-map/)
  makes the current ownership, reader paths, two compatible renderings, and
  missing bridges inspectable in one place. It is a local routing prototype,
  not the canonical public hub or evidence index.
- The Vizier assessment remains a separate application-specific companion. It
  can consume the field synthesis and experiment agenda without turning the
  state-of-the-field report into a product evaluation.

## What the current package already does well

The two long reports and executive summary already provide a strong core:

- a literal definition of the field and a typology of visualization purposes;
- the distinction between possible output, accomplished work, and creator or
  reader experience;
- a 2023–2026 capability and benchmark timeline without inventing one universal
  progress curve;
- a context-sensitive ideal and a fixed-scale 2017–2026 capability scorecard
  that keeps maturity, evidence confidence, and forecast probability separate;
- a synthesis of convergent techniques, legitimate divergences, and direct
  negative or null findings;
- an evidence ladder from structure and execution through reader outcome;
- a job taxonomy, work-surface taxonomy, practitioner adoption evidence, and an
  evidence-graded map of who appears to reach for what;
- a clear separation between creator satisfaction, artifact correctness, and
  reader comprehension; and
- gap ledgers that say what is partly answered, what is merely demonstrated to
  fail, and what remains open.

Those are the central intellectual assets. The packaging problem is now less
about adding another framework than connecting these assets into one argument
and filling the places where the title still promises more than the evidence
package delivers.

## What is missing or underbaked

| Area | Current state | Why it matters | Needed next |
| --- | --- | --- | --- |
| **Capability-to-practice bridge** | The benchmark crosswalk now makes the first half explicit: each measured capability is paired with its input, output, grader, human context, supported claim, and nonclaim. The chain from that bounded capability to changed work, accepted delivery, and reader consequence remains distributed. | Readers need to see why a benchmark gain does—or does not—change ordinary work. | Extend the crosswalk into a chapter mapping each capability and technique to the job, work surface, acceptance evidence, and downstream reader consequence it can affect. |
| **Complete system anatomy** | Individual mechanisms are well described; the end-to-end architecture is distributed across sections. | “Use a critic” or “use a semantic layer” is meaningless without knowing what information and authority each component owns. | One reference architecture showing intent, data contract, representation, generation, execution, critique, delivery, and human gates, with environment-specific substitutions. |
| **Tool landscape versus adoption** | The tool-family showcase is strong; named-product adoption remains mostly unknown. | Availability, provider target, incumbent co-use, stars, installs, and routine use are different signals. | A durable landscape record with dated feature contracts and separately labeled adoption proxies; no synthetic market-share ranking. |
| **Agent skill packages** | The deep dive now compares 18 files or families, registry installations, repository activity, broader skill benchmarks, and one paired scientific-visualization study. | Skills may encode genuinely useful checks, duplicate the harness, or add obsolete and conflicting doctrine. | Independent current-model tests of the popular generic, dashboard, accessibility, mobile, and explanatory packages; maintenance tests after model and library releases. |
| **Specialized vision and critique models** | The deep dive now separates chart QA, parsing, OCR/layout, grounding, integrity or quality critique, and verifier/repair roles. It finds component-level specialist gains and newer transfer tests where general models lead. | A specialist may be valuable for exact extraction or defect detection even when it is not the best general reasoner. | Independent end-to-end routing tests at equal budget, including source fidelity, interaction, mobile delivery, repair success, cost, and reader outcomes. |
| **Production lifecycle** | Interviews with 17 biomedical-visualization practitioners show that high-stakes teams often restrict AI to auxiliary work; controlled studies still stop before authenticated delivery, handoff, later correction, maintenance, or retirement. | First-render quality is a poor estimate of production value and risk. | Production cases or original tests that follow accepted artifacts through deployment, later data/model changes, regression, and ownership transfer. |
| **Economics and abandonment** | An 18-analyst forced-error study records unfinished work and premature acceptance without a detected time advantage among its interface conditions; other studies report fragments of iteration or inference budget. | Teams choose workflows on total human and machine cost, not benchmark score alone. | A common ledger for prompting, waiting, verification, failed repair, model spend, deployment, abandoned work, and maintenance. |
| **Reader outcomes, accessibility, and mobile** | Controlled evidence now includes harm from selected AI-generated misleading charts and strong adjacent evidence that low-vision and nonvisual outcomes depend on the actual mobile or assistive interaction. No study joins AI generation to that delivered-reader evidence. | The final audience experiences the delivered surface, not the authoring session. | Reader studies with declared literacy, task, device, assistive technology, comprehension, confidence, decision, and harm measures. |
| **Human capability and learning** | The human-skills review separates consumption, construction, critique, and connection from data, domain, tool, situated-judgment, delivery, and AI-interaction resources. The learning chapter adds a three-period development arc, a practice portfolio, one positive immediate post-removal result for proactive scaffolding, adjacent randomized evidence of performance-learning dissociation, and visualization-specific ordinary retention evidence. | “Novice” and “expert” remain inconsistently defined. Delayed visualization construction and far transfer remain mostly unmeasured, and no captured longitudinal study causally estimates AI-driven atrophy. | Capability-profiled trials crossing answer-oriented and metacognitive assistance with immediate withdrawal, delayed retention, unfamiliar transfer, verification, and production-to-reader follow-through. |
| **Environment case evidence** | The typology explains why journalism, BI, operations, science, exploration, and education differ; direct evidence remains uneven. | A universal recommendation becomes wrong when purpose, stakes, update cycle, and publication authority change. | One evidence-bearing case chapter per major environment, including a representative task, failure boundary, and acceptance test. |
| **Organizational adoption and governance** | Semantic models and provider controls are described; organizational ownership, procurement, review, and incident response are thin. | Governed data does not govern itself, and one wrong executive number can terminate trust. | Field evidence on ownership, refusal behavior, review queues, audit, escalation, and recovery after error. |
| **Security, privacy, provenance, and disclosure** | These appear as boundaries but not as a coherent state-of-the-field chapter. | Data access, generated code, external services, source lineage, and AI disclosure can determine whether a workflow is usable at all. | A scoped chapter separating security controls, data authority, analytical provenance, publication disclosure, and reader-facing source evidence. |
| **International and non-English use** | Multilingual POLYCHARTQA now measures chart QA across ten languages, and MM-JudgeBench measures multilingual judge behavior across 25 languages with a chart-centric subset. Both reveal language-dependent failures but derive from translated English-centric material. | Language, chart convention, infrastructure, and institutional context may change prompts, authoring, and reader interpretation. | Human-authored tasks and delivered-reader studies sampled from ordinary non-English work, including local chart conventions, low-resource languages, code-switching, and mobile contexts. |
| **Living update mechanics and forecast calibration** | The reports have an evidence cut and refresh triggers; the forecast adds seven dated bets with stable IDs, explicit resolution tests, confidence, annulment conditions, and signposts. The capability scorecard now supplies a public fixed rubric, retrospective scored history, and conditional future score movements. A public change log and recurring calibration record remain missing. | Models, harnesses, products, registries, and benchmarks change faster than a monolithic report can be rewritten, and an unscored forecast can become persuasive prose. | Dated change notes; quarterly signpost review; resolution and calibration when forecasts come due; a rule for updating an essay without silently rewriting the historical snapshot. |
| **Original comparison evidence** | The synthesis is stronger than the local experiment record. | Several decisive questions—skill lift, critic routing, context routing, total cost—cannot be settled by more literature review. | A small sequence of preregistered, equal-budget experiments using the current model plus normal harness as the baseline. |

## Comprehensive table of contents

Every numbered chapter below can also be a standalone essay. The comprehensive
rendering adds introductions and transitions at the part level; the essay
rendering adds a local definition, evidence note, and “where this fits” link.

### Front matter

0. **Executive summary: the state of AI-assisted data visualization in August
   2026**<br>
   The conclusions that survive across research systems, products, practitioner
   evidence, and reader studies; what is less settled; the five decisions a
   reader can make from the report.

1. **Scope, vocabulary, and how to read the evidence**<br>
   What counts as AI-assisted visualization; agent, harness, skill, semantic
   model, critic, benchmark, ablation, and held-out task; evidence grades and
   source vintages.

### Part I — What the field is trying to accomplish

2. **The goals of AI-assisted visualization**<br>
   Established task typologies; explanation, exploration, monitoring, decision,
   science, education, and production as different objectives.

3. **From possible output to accomplished work to reader experience**<br>
   The complete creator-to-artifact-to-reader episode and why lower thresholds
   do not stand in for higher ones.

4. **The actors and jobs**<br>
   The twelve-job workflow; creators, data owners, reviewers, publishers,
   maintainers, readers, and decision-makers; where responsibility remains.

### Part II — What systems can do

5. **The ideal state and the same capability scorecard through time**<br>
   A shared nine-dimension spine; context-specific definitions of perfect;
   current maturity and confidence; the 2017–2026 reconstruction on one fixed
   rubric; the model, harness, system, and benchmark milestones that justify
   threshold changes; and what can and cannot be projected.

6. **Chart and dashboard generation**<br>
   Static charts, chart-to-code, multi-view dashboards, scientific visualization,
   interaction, and the gap between execution and semantic success.

7. **Understanding data, charts, and visual integrity**<br>
   Data extraction, chart question answering, literacy, misleading-chart
   detection, semantic reasoning, and why these are separate abilities.

8. **Specialized vision models for parsing and critique**<br>
   Chart-specific models, OCR and table recovery, visual grounding, perceptual
   critique, verifier roles, same-benchmark comparisons, and routing criteria.

9. **What benchmarks actually measure**<br>
   A benchmark crosswalk from code/spec similarity through execution, fidelity,
   integrity, interaction, analytical support, and reader outcome.

### Part III — Techniques and system designs

10. **The technique spine**<br>
    Typed representations, deterministic execution, semantic grounding,
    mixed-initiative controls, visual inspection, browser replay, provenance,
    and human gates.

11. **Representations, DSLs, and compilers**<br>
    Where constrained generation wins, where direct code remains appropriate,
    and the trade between portability and expressiveness.

12. **Grounding in the real data contract**<br>
    Schema, measures, units, joins, verified queries, permissions, examples,
    source vintage, and the limits of semantic layers.

13. **Critics, repair loops, and stopping rules**<br>
    Text, code, numeric, and visual feedback; fresh context; regressive editing;
    critic knowledge; diminishing returns; equal-budget evaluation.

14. **Single agents, multiple agents, and tool boundaries**<br>
    When specialization adds information or executable capability, and when
    persona-labeled roles only add inference and coordination cost.

15. **Visualization skills inside coding agents**<br>
    Registry landscape, package anatomy, convergent and conflicting guidance,
    maturity ladder, adoption proxies, negative transfer, and decay after model
    releases.

16. **Mixed-initiative authoring and human control**<br>
    Natural language beside direct manipulation, visible intermediate data,
    branches, history, local edits, reversion, and expertise-sensitive support.

### Part IV — What people use and experience

17. **Who is reaching for what**<br>
    Practitioner adoption evidence, audience-to-work-surface map, incumbent
    co-use, provider targets, and the boundary between observation and inference.

18. **The current tool landscape**<br>
    General assistants, spreadsheets, notebooks, guided authoring, governed BI,
    conversational consumption, and open/developer systems—compared by what
    they see, change, expose, deliver, and govern.

19. **The creator experience**<br>
    First-draft speed, alternatives, repetitive work, correction loops,
    verification, confidence, satisfaction, opportunity cost, and abandonment.

20. **Expertise changes both leverage and risk**<br>
    Consumption, construction, critique, and connection; data, domain, tool,
    judgment, delivery, and AI-interaction resources; novice access and error
    detection; learner confidence versus transfer; expert filtering and
    opportunity cost; bounded gains and the absence of skill-atrophy evidence.

20A. **Learning visualization when generation is cheap**<br>
    Recent-past, present, and forecast learning bottlenecks; what becomes easier
    to do versus easier to retain; productive difficulty; practice to invest,
    maintain, or de-emphasize; scaffold withdrawal; delayed transfer; and the
    learner-to-reader path.

21. **The reader experience**<br>
    Comprehension, trust cues, disclosure, provenance, visual literacy,
    decisions, uncertainty, accessibility, mobile reading, and why model chart
    reading is not a reader test.

22. **How context changes the answer**<br>
    A master comparison across journalism, operational dashboards, analytical
    BI, exploration, science, education, public communication, and reusable data
    applications. Each environment can also become its own case essay.

### Part V — What the evidence says to do

23. **Where the field is converging**<br>
    Inspectable intent, real data contracts, deterministic evidence, mixed
    initiative, delivered-artifact evaluation, and recoverable human control.

24. **Where approaches legitimately diverge**<br>
    Code versus DSL, prompt-first versus direct control, open files versus
    governed systems, universal versus local contracts, static versus
    interaction-aware evaluation.

25. **What performs poorly, adds no value, or remains unsupported**<br>
    Direct negative and null findings; generic personas; unlimited critic loops;
    prompt examples without lift; synthetic readers; nonblank renders; stars and
    installs as outcome evidence.

26. **The two gap ledgers**<br>
    Research capability and technique gaps beside practitioner and reader gaps;
    partly answered questions, demonstrated failures, open measurements, and
    refresh triggers.

27. **How to evaluate a real workflow**<br>
    Representative tasks, current baseline, first draft versus acceptance,
    consequential choices, delivered-state testing, reader testing, total cost,
    and evidence retention.

28. **Research and product experiment agenda**<br>
    Equal-budget critic comparison; skill ablation; context routing; semantic
    fidelity; longitudinal delivery and maintenance; mobile/accessibility reader
    study; stopping and promotion rules.

29. **What would unlock the next capabilities?**<br>
    Capability gates by environment; raw-model versus technique-led progress;
    dated forecasts for static work, interaction, critics, specialists,
    production, readers, and autonomy; signposts, falsifiers, and the local
    context, authority, provenance, and acceptance evidence that scale alone
    cannot supply.

### Appendices

A. **System and tool catalog** — current feature contracts, version dates, and
evidence class.<br>
B. **Benchmark crosswalk** — tasks, data, outputs, metrics, models, judges,
release status, and comparability limits.<br>
C. **Study cards** — every human study explained in plain language with sample,
task, measures, result, and non-result.<br>
D. **Evidence and source index** — primary sources, pinned repositories, capture
dates, grades, and refresh conditions.<br>
E. **Glossary** — one definition for every recurring specialist term.<br>
F. **Change log** — what changed after the August 2026 snapshot and why.

## A practical essay series

The master contents yields four useful entry routes rather than one publication
sequence:

- **Start with capability:** chapters 5–9, then 23–26.
- **Start with building systems:** chapters 10–16, then 27–29.
- **Start with practice:** chapters 3–4 and 17–22, then 26–27.
- **Start with one contested claim:** chapter 25, followed by the relevant
  benchmark, technique, or experience chapter.

Each essay should contain five local elements:

1. the question it answers;
2. a literal short answer;
3. the minimum vocabulary needed for that answer;
4. the strongest supporting, conflicting, and missing evidence; and
5. links to the preceding and following concepts in the comprehensive report.

No post should depend on a backreference for its central distinction. No number
should appear because a summary card looks empty without it. No product grid
should imply comparative quality unless the comparison actually exists. The
shared evidence base—not duplicated prose—is the product that allows both
renderings to stay coherent.

## Update log

- **2026-08-14 — Initial public edition.** Defined the comprehensive report and
  non-linear essay renderings, content ownership, shared evidence spine,
  complete table of contents, and underbuilt bridges.
