# Executive summary: AI-assisted data visualization in August 2026

**Status:** state-of-the-field snapshot. **Evidence cut:** 2026-08-14. **Recheck
by:** 2026-11-14, or after a material model, harness, benchmark, registry, or
analytics-assistant release.

This summary draws on the [research synthesis](/reports/research-review/),
the [practitioner and reader report](/reports/practitioner-and-reader-experience/),
the [human-skills and banked-gains review](/reports/human-skills-and-gains/),
the [learning-with-AI chapter](/reports/learning-with-ai/),
the [agent-skill deep dive](/reports/agent-skills/), and
the [specialized-vision deep dive](/reports/specialized-vision-models/).
The [capability forecast](/reports/capability-forecast/)
states which gains would unlock the remaining levels, separates model and
technique effects, and records seven dated resolution tests through August
2029.
The [editorial architecture](/reports/editorial-architecture/)
shows how these can become one comprehensive report and a set of self-contained
essays without duplicating the evidence base.

**Visual entry points:** explore the [research field
guide](/research/),
the [current and historical capability
scorecards](/capability-scorecard/),
the [practitioner and reader
experience](/practitioner-experience/),
or the [map of the complete research
package](/content-map/).

## The short version

**AI can now produce useful charts, chart code, scientific views, and
increasingly ambitious interactive dashboards.** It can accelerate first
drafts, translate between unfamiliar representations, automate repetitive
implementation, and make visualization accessible to people who could not
otherwise build one.

**That is not the same as reliably producing an accepted visualization.** The
remaining failures are often the ones that matter most: hidden transformations,
wrong business meaning, regressive corrections, interaction state, mobile and
assistive delivery, maintenance, and whether a defined reader understands the
claim. **A fluent chart can hide these failures rather than expose them.**

The field is not converging on one agent, product, or universal workflow. It is
converging on a set of controls:

1. **Declare context:** purpose, audience, data contract, delivery surface, and
   stakes;
2. **Expose the work:** transformations, encodings, state, and revision
   history;
3. **Compute what can be computed:** use deterministic methods for claims that
   can be checked deterministically;
4. **Test the delivered artifact:** inspect and exercise the actual output; and
5. **Keep consequential acceptance human:** reserve contextual, consequential,
   and reader-level acceptance for people with the relevant authority or
   experience.

**Natural language is becoming one control surface inside a structured
environment—not a replacement for the environment.**

## Capability improved quickly, but the target also became harder

From 2023 through August 2026, general models gained image understanding, code
execution, tool use, browser operation, persistent workspaces, and reusable
skills. Research advanced from single-chart generation and chart-to-code tasks
to editable multi-view dashboards, replayable interaction, realistic chart
reading, visual integrity, and expert-adjudicated critique.

Within individual benchmarks, later systems often make **large gains in visual
fidelity and task breadth**. Execution and semantic reliability improve less
uniformly. **Across benchmarks, there is no honest single progress curve**
because the inputs, models, renderers, judges, and definitions of success change.
Stronger baselines also expose harder questions: not merely whether the chart
renders, but whether the data binding is right, the interaction works, a repair
introduces a regression, or a reader learns the intended point.

Two August benchmarks sharpen that boundary. **Dashboard2Code** now tests stateful
reconstruction across 180 Plotly Dash applications and 450 interaction tasks;
hidden state and transformations can be wrong while the interface responds.
**Chartography's best tested configuration reaches 45% mean pass@1** on 100
deliberately difficult, practitioner-authored professional chart-reading tasks.
The latter is not an ordinary-chart failure rate. Together they show that
interactive state and professional visual conventions are distinct capability
frontiers, not details covered by a generic screenshot score.

The [benchmark crosswalk](/reports/benchmark-crosswalk/)
shows exactly what each major evaluation measures and cannot establish. The
[capability scorecard](/reports/ideal-state-and-scorecards/)
turns that evidence into current, historical, and forecast scores by use context.

The forecast derived from this record is deliberately asymmetric. It assigns
higher near-term probabilities to accepted static work, interactive-dashboard
reasoning, narrow visual critics, and environment-specific adapters because
they have public tests and observable failure signals. It assigns lower
probabilities to production productivity, representative-reader benefit, and
autonomous publication because the evidence needed to verify those outcomes is
mostly absent. Better models can help request and reason over local definitions,
provenance, authority, and reader response; they cannot manufacture those facts.

## What consistently helps

- **Real context.** Field definitions, measures, joins, units, permissions,
  source vintage, audience, and delivery constraints prevent errors a generic
  persona cannot.
- **Inspectable representations.** A semantic chart object, analytic
  specification, restricted language, or visible intermediate table makes
  validation and local correction easier. Direct code remains appropriate
  where expressiveness matters and the risks are low.
- **Mixed-initiative control.** Natural language works better beside direct
  manipulation, visible transformed data, history, branching, and reversion
  than as an endless regenerate button.
- **Layered verification.** Query execution, recalculation, schema checks,
  rendered inspection, browser replay, accessibility checks, and human review
  establish different facts. None substitutes for the others.
- **Bounded specialization.** A separate agent, skill, critic, or vision model
  helps when it owns different information, a tool, or an evidence channel.
  A new job title alone does not create specialization.

**These are shared principles, not a single architecture.** Journalism, governed
BI, operational monitoring, exploration, science, education, and reusable
applications have different purposes, authorities, update cycles, and failure
costs. A common record of intent, data, transformations, artifact, revisions,
and evidence should route to different authoring and acceptance procedures.

## What creators and readers actually experience

The [practitioner and reader experience](/practitioner-experience/)
explains why enthusiasm and frustration coexist. **First drafts can appear
before an idea has cooled, and tedious implementation can compress
dramatically. Correction and verification can erase that gain.**
In one study of 18 experienced analysts completing 108 deliberately error-prone
tasks, **seven episodes were unfinished and participants prematurely accepted
the result 31 times**. A novice study found low verification intent and repeated
failed repairs alongside positive satisfaction.

Expertise changes both the benefit and the burden. Experienced practitioners
can filter weak suggestions and use AI selectively, but they also notice when
describing a small edit takes longer than making it. In interviews with 17
biomedical-visualization practitioners, AI was used chiefly for auxiliary work;
all opposed substantial generated visuals in final scientific communication.

Expertise is better represented as a capability profile than a rank. Established
visualization-literacy research distinguishes consuming, constructing,
critiquing, and connecting a chart to its context; professional work adds data,
domain, implementation, situated-judgment, and delivery resources. Current
learner studies bank access, reported speed, confidence, and engagement more
readily than correctness, verification, or delayed independent skill. One
117-person randomized visual-comprehension study found an immediate
post-removal advantage for a proactive agent that asked scaffolded questions;
it did not test delayed construction or far transfer. Adjacent randomized
education studies show why the distinction matters: assisted exercise
performance can improve while conceptual learning does not, and unrestricted
answer access can harm later unassisted work. Small studies suggest intermediate
and expert practitioners can turn structured
critique and constrained implementation into better work more reliably, but no
large study establishes one universal expertise curve.

**Readers receive the claim, not the authoring transcript.** In a controlled
48-person study, selected AI-generated misleading charts reduced answer
accuracy **from 88.3% to 71.9%**. Adjacent accessibility studies show why delivery
cannot be inferred from a responsive screenshot or text alternative: outcomes
for low-vision and blind readers changed with the actual device, interaction,
modality, chart type, time, and workload. A new 12-person blind and low-vision
study adds direct AI-assisted learning evidence: eleven preferred tactile
charts plus text and chat and described a stronger spatial model, but measured
chart-understanding accuracy did not improve. No current study joins AI
generation to a representative mobile or assistive-technology reader
evaluation.

Following the artifact past its creator exposes another class of failure. In
one public second-maintainer episode, an AI-built dashboard failed during the
creator's absence because refresh and scheduling lived on that person's
laptop; the inheritor replaced it with a proper pipeline. This is testimony,
not a prevalence estimate, but it makes execution host, lineage, definitions,
dependencies, documentation, ownership, and rollback part of acceptance—not
optional handoff cleanup.

## Skills can help; popularity does not say whether they do

The [agent-skill deep dive](/reports/agent-skills/)
finds that the public ecosystem contains hundreds of listings labeled as data
visualization, with extensive copying, vendoring, and source drift. Registry
installs measure acquisition. Repository stars usually belong to a much larger
repository. **Neither measures routine use or output quality.**

The inspected packages perform six different jobs: primers, storytelling
guides, renderer gateways, library or domain adapters, quality gates, and
structured workflows. Their most defensible value is information a capable
model cannot reliably infer: a current API, local environment, semantic
representation, known failure, executable validator, or delivered-surface
repair loop.

There is direct evidence that this can work. **SciVisAgentSkills improved quality
in all ten paired suite-by-agent comparisons** across 108 scientific-
visualization tasks, although one completion measure fell. Broader benchmarks
show the boundary: compact, relevant, compatible skills can help; comprehensive,
self-generated, stale, or incorrectly retrieved guidance can perform worse than
no skill. Popular generic, dashboard, accessibility, mobile, and explanatory
packages still lack independent paired evaluation.

## Specialized vision is a sensor, not a final judge

The [specialized-vision deep dive](/reports/specialized-vision-models/)
finds that “vision for charts” includes at least six jobs: question answering,
structured extraction, OCR and layout, element grounding, integrity or
perceptual critique, and verification or repair. No model covers all six
reliably.

Specialists show real component gains. VisJudge-7B matched its expert-
adjudicated quality rubric better than the tested general models. ChartAgent's
chart-specific tools materially improved the same general reasoner on numeric
questions. Chart-aware grounding and recent OCR/layout models improve bounded
perception. At the same time, newer out-of-distribution extraction and realistic
chart-QA tests show strong general models overtaking older chart specialists.
Dashboard2Code adds a stateful fixed-desktop test but leaves responsive/mobile,
animation, keyboard, and assistive-technology behavior open. Chartography shows
that domain conventions and hard professional visual forms still defeat strong
general systems even with more reasoning.

**The practical choice is role-based:** use source data and deterministic
checks whenever available; add a specialist for a measured parsing,
localization, or perceptual failure; retain a current general model for broad
reasoning; and do not confuse any model's score with reader comprehension or
publication acceptance.

## What performs poorly, adds risk, or remains unsupported

- **Proxy as proof:** treating a nonblank render, attractive screenshot,
  code-similarity score, or one model judge as proof of correctness;
- **Unbounded repair:** unlimited self-critique or repair without a bounded
  fault, budget, regression check, and stopping rule;
- **Role-play instead of specialization:** multiplying agents or personas
  without distinct information and tool boundaries;
- **Synthetic readers:** using personas as representative readers or domain
  experts;
- **Popularity as evidence:** assuming a skill helps because it is long,
  popular, installed, or official;
- **Specialists outside their lane:** applying a chart specialist outside the
  distribution and role it was tested on;
- **Premature storytelling:** forcing a narrative arc before the analysis
  supports one;
- **Desktop shrinkage:** shrinking a desktop composition and calling the result
  mobile; and
- **First-render accounting:** claiming production value from first-render
  speed without counting prompting, waiting, verification, failed repair,
  deployment, abandonment, and later maintenance.

## The next evidence should follow work to acceptance

The [capability forecast](/reports/capability-forecast/)
and [research-package map](/content-map/)
show where evidence is still missing. **The largest gaps sit between measured
capability and lived outcome:**

- **Acceptance:** the share of ordinary work accepted without repair;
- **Semantics:** fidelity against owned measures and changing sources;
- **Correction:** cost, regression, stopping, and abandonment;
- **Routing:** equal-budget comparisons among deterministic checks, general
  critics, specialist models, and skills;
- **Who gains and how:** capability-profiled gains across access, productivity,
  quality, learning, verification, reader outcome, and sustained work;
- **Learning:** delayed unassisted tests that distinguish learning from
  dependence or skill atrophy;
- **Operations:** authenticated delivery, multi-context handoff, maintenance,
  and total cost;
- **Governance:** organizational ownership, privacy, provenance, disclosure,
  and recovery after error; and
- **Reader outcomes:** comprehension, decisions, calibrated trust,
  accessibility, and mobile use by the intended readers.

**The next experiments should begin with the current model and its normal harness,**
use representative tasks from declared environments, separate first draft from
accepted artifact, add one mechanism at a time, exercise the delivered surface,
and retain every correction and regression. Additional instructions, agents,
or models should survive only when they improve that full path for a stated job
at a proportionate total cost.

## Update log

- **2026-08-14 — Readability and navigation update.** Added selective emphasis
  to the load-bearing conclusions and linked claims to the relevant visual
  experiences and deep-dive reports.
- **2026-08-14 — Initial public edition.** Condensed the research, practitioner,
  human-skills, learning, agent-skill, specialized-vision, and forecasting work
  into one decision-oriented summary.
