The state of AI-assisted data visualization Markdown source

Research snapshot · Evidence reviewed through August 14, 2026

Human skills and banked gains in AI-assisted data visualization

Status: research snapshot, evidence cut 14 August 2026. This chapter examines what a person needs to know or be able to do when making and reading data visualizations, then asks which of those capabilities AI has demonstrably improved for different people. It is a companion to The state of AI-assisted data visualization research and The practitioner and reader experience. The companion chapter Learning data visualization when AI can make the chart follows the same capabilities across the recent past, August 2026, and the next three years, then turns the evidence into a practice allocation and learning loop.

Executive summary

Visualization expertise is not one thing, and “novice” is not one population. A person may read familiar business charts fluently but be unable to program; know a scientific domain but not visualization design; or build polished charts without being able to audit a misleading transformation. Studies that label all of these people novices or experts conceal the mechanism that AI is helping.

The most useful existing framework defines visualization literacy through four families of competence: consume a visualization, construct one, critique its evidence and design, and connect it to a domain, audience, decision, or public context. Professional work also draws on data and statistical knowledge, domain semantics, implementation skill, situated design judgment, and collaboration or delivery knowledge. AI can reduce the cost of some of these resources. It does not make the whole profile uniform.

The evidence supports five conclusions.

  1. The best-banked gain for less experienced people is access, not autonomy. Natural language lets people attempt transformations, code, chart forms, and critique that were previously out of reach. In controlled education studies, AI use also reduced reported time, expanded idea lists, and increased confidence or engagement. These studies found much weaker evidence of durable visualization proficiency, independent transfer, or error detection.
  2. Intermediate and experienced practitioners appear to convert assistance into better work more reliably, but the evidence is small and bounded. In a 13-designer critique study, intermediate and expert designers used advice more effectively than novices. In a seven-expert scientific-visualization study, a restricted language and compiler reduced debugging and improved bounded replication work. These are promising mechanisms, not population estimates.
  3. AI substitutes for syntax and search more readily than for situated judgment. Professional designers combine framing, quality, appearance, tool, navigation, and composition judgments in ways shaped by the client, data, medium, precedent, and moment. Experienced analysts similarly describe exploration as opportunistic and domain-guided. Generic rules or a prompt cannot fully specify that work in advance.
  4. Verification is a human capability and a system property. People verify through representations they understand: source values, intermediate tables, SQL, Python, visual output, or direct controls. A current novice study found only three data-verification attempts across 60 sessions. A study of experienced analysts still recorded premature acceptance in 31 of 108 deliberately error-prone episodes. Expertise helps, but does not make an opaque or misleading workflow safe.
  5. There is one encouraging immediate learning result, but not yet convincing evidence of durable skill atrophy—or durable AI-taught construction skill. A 117-person randomized visual-comprehension study found that a proactive agent using scaffolded questions outperformed a passive agent and a data story after assistance was removed. The test was immediate and used the same visualization formats. Elsewhere, passive acceptance, reduced exploration, and dependence are visible risks, but delayed unassisted transfer remains rare. Claims that AI either teaches visualization broadly or destroys the skill still run ahead of the evidence.

The practical design rule is to measure a specific gain against a specific capability profile. Do not ask whether “novices benefit.” Ask whether, for example, a domain expert with low coding fluency can produce a correct, inspectable analysis faster; whether a design student can later critique a new chart without AI; or whether an experienced author can incorporate useful feedback without spending more time filtering it than editing directly.

A better model of visualization skill

The 2025 state-of-the-art review of visualization literacy reviewed 374 papers and proposed four overlapping competency themes. It is the best available standard spine because it includes more than chart reading and requires every competency to be operationalized in a domain, scenario, and audience.

Competency Literal question Examples What AI may change
Consumption Can the person read and interpret the representation? locate and compare values; read axes and legends; recognize patterns; interpret uncertainty explain unfamiliar forms, answer questions, direct attention, translate a chart into prose or another representation
Construction Can the person produce an effective representation? select data; choose a form; map fields to marks and channels; implement, annotate, and revise remove syntax barriers, propose encodings, write transformations and code, generate variants, apply repetitive changes
Critique Can the person judge the chart and its production process? identify misleading scales or aggregations; inspect source fidelity; question omissions; assess whether the design supports the claim enumerate checks, run tests, inspect renders, surface common design failures—but also produce plausible rationalizations and false reassurance
Connection Can the person relate the visualization to external meaning and action? use domain knowledge; integrate a chart into a story; debate implications; make a decision; recognize who or what is missing retrieve background, bridge creator and domain language, draft a narrative—without owning local definitions, consequences, or publication authority

These are competency themes, not four stages and not four kinds of people. A journalist may need all four to integrate evidence into a story. An operational analyst may construct only familiar dashboards but make sophisticated connections to a live business process. A reader may need no implementation skill and still need strong critique and connection skills to judge a public claim.

The resources behind the four competencies

The four-part spine becomes more diagnostic when it is crossed with the resources that real work draws upon.

Human resource What it contributes What happens when it is missing
Data and statistical knowledge aggregation, comparison, missingness, uncertainty, method choice, and causal restraint a well-drawn chart can answer the wrong quantitative question
Domain and semantic knowledge field meaning, measure definitions, plausible ranges, institutional history, and decision consequences the output can be structurally valid and locally false
Visual and perceptual knowledge encoding, hierarchy, comparison, annotation, accessibility, and composition the result can be correct but hard to read, misleading, or poorly matched to its audience
Tool and implementation knowledge transformation, code, direct manipulation, interaction, responsive behavior, and production constraints ideas remain unbuildable, uninspectable, brittle, or expensive to revise
Critical and epistemic judgment verification, source assessment, alternative explanations, stopping rules, and calibrated trust speed and polish become substitutes for evidence
Situated design judgment framing, quality, appearance, tool choice, navigation, and composition in the specific project generic best practice displaces the actual purpose, material, client, or medium
Collaboration, audience, and delivery knowledge eliciting requirements, handoff, governance, accessibility, publication, maintenance, and reader testing an acceptable draft fails as shared or delivered work
AI interaction literacy scoping, specifying intent, exposing assumptions, comparing alternatives, inspecting state, correcting locally, and refusing or rolling back the person either underuses a useful system or accepts a fluent but ungrounded result

The first seven resources predate generative AI. The eighth is new, but should not be confused with prompt ornament. It is primarily the ability to establish an explicit work contract and evaluate what the system did.

“Novice” and “expert” are unsafe shortcuts

Burns and colleagues analyzed 79 influential visualization papers that referred to novices, non-experts, laypeople, or the general public. They found disagreement within and between papers about what defined a novice, plus a mismatch between broad claims and study samples concentrated among young people, students, and US residents. Papers variously used profession, scientific proximity, technical skill, visualization experience, and the absence of some knowledge as the boundary.

That problem is acute in AI research. A model can remove a programming barrier while leaving visualization, data, and domain barriers untouched. It can also help an experienced coder who is a novice in the dataset, or a subject-matter expert who has never built an interactive chart. A single expertise label makes these distinct gains look equivalent.

A defensible study should therefore publish a capability profile rather than a rank alone:

The label can remain as a compact description after those dimensions are visible. It cannot replace them.

What expertise looks like in observed work

Three pre-AI lines of research clarify what assistance is trying to augment.

The implication for AI is not that guidance is useless. It is that a universal recipe can support elementary construction while conflicting with the situated, layered judgment of advanced practice.

What counts as a banked gain

The phrase banked gain is intentionally stricter than “the participant liked the system” or “the model produced a chart.” A gain is banked when it survives at the outcome level that matters for the task.

Gain Evidence that counts What does not establish it
Access the person successfully attempts a task or representation they could not complete under the declared baseline a nonblank render or stated empowerment
Productivity less total human and machine time or effort to an accepted artifact of comparable quality time to first draft, self-reported ease, or omitted verification time
Artifact quality a better accepted output under a declared correctness, usefulness, design, or delivery rubric polish, one model judge, or satisfaction alone
Learning better unassisted post-test, delayed retention, or transfer to a new task confidence, engagement, AI-assisted final work, or prompt skill alone
Judgment and verification more defects detected and correctly repaired, better calibration, or fewer false acceptances more explanations or perceived control
Reader outcome better comprehension, recall, decision, trust calibration, accessibility, or appropriate action creator success or model chart-question answering
Sustained work maintainability, later correction, handoff, reuse, or regression performance one session or a published screenshot

A system can bank one gain while losing another. Faster completion with less diverse exploration is a productivity gain and a possible learning or creativity cost. Higher confidence paired with pervasive errors is an experience gain and a calibration failure. The outcomes should not be averaged into one “AI helps” score.

Where gains have actually been observed

The table is organized by people and outcomes, not by product. “Measured” means the study observed behavior or assessed artifacts; it does not mean the result generalizes beyond that study.

Population and study Assistance Gain that was observed What was not banked
95 masters students in a group data-story exercise (Ahmad & Ma, 2024) general LLMs, including a 30-minute prompting lesson the non-LLM groups reported spending 44% more time; AI users reported more focus on insights; groups with prior LLM experience received higher scores in a secondary analysis most AI versus non-AI artifact differences were not statistically significant; non-AI groups scored slightly higher on creativity; non-use was not verified; time was self-reported
30 graduate students over three sequential workshops (More Than Chatting) peer learning, then required LLM use and prompt instruction, then optional use confidence, perceived utility, and engagement increased; pre/post competency effect was small to moderate (d=.358); expert scores improved modestly over the sequence no parallel control separated AI from instruction, practice, task, or workshop order; the paper itself reports a gap between engagement and skill
117 higher-education participants in a randomized comprehension experiment (Yan et al., 2024) data story, passive conversational agent, or proactive agent using scaffolded questions all conditions improved; the proactive agent produced better immediate comprehension after assistance was removed same visualization formats at post-test; immediate rather than delayed transfer; comprehension rather than construction
26 students across four semester projects (Kim et al., 2024) general conversational assistant used with Tableau, D3, and Vega-Lite work reported speed, confidence, coding help, debugging, and access outside office hours; 3,773 queries document actual use no parallel control or delayed transfer test; query themes, volume, and length were not significantly related to grades; design help was less central
30 graduate students in a complementary encoding exercise (Ahmad & Ma, 2024) LLM encoding suggestions versus students’ own knowledge more idea and phrasing options were reported only 14 completed the post-survey; students chose less diverse visual encodings when given AI suggestions; perceived creativity was neutral
20 visualization novices across 60 sessions (Vibe Visualizing) current general assistant used for authoring and interpretation access, confidence, and satisfaction: participants produced artifacts and rated mean confidence 3.73/5 and satisfaction 3.93/5 all charts had flaws; data verification appeared only three times; 11/16 clutter-removal repairs and 8/9 unusable-chart repairs failed; confidence was not calibrated to quality
13 designers: six novice, four intermediate, three expert (Visualizationary) LLM and deterministic perceptual critique, hierarchy, and version history on self-selected work external experts rated final change 3.69/5 on average; intermediate and expert means were higher than the novice mean; 68 critiques were used selectively no no-critique baseline; roughly 90–150 observed minutes per participant; some participants did not end on their best-rated version; no production or maintenance outcome
12 practitioners bringing their own work (Kim et al., 2025) GPT-3.5 advice compared with seasoned visualization experts rapid brainstorming, broad option generation, and a neutral first pass practitioners preferred human experts for accuracy, helpfulness, reliability, context, adaptability, and actionable conversation; raw model capability is dated
18 experienced analysts across 108 forced-error episodes (Shang et al., 2024) conversational, stepwise, or phasewise AI analysis interfaces structured decomposition improved perceived control no detected difference in task success, time, or verification hints; seven episodes were unfinished and 31 were accepted with an issue remaining
Seven visualization experts on three scientific replication tasks (Raiven, 2026) restricted visualization language, deterministic compiler, and visual feedback bounded ease and efficiency in reproducing specified scientific views; five preferred the system and two had no preference; compiler-based generation sharply reduced execution failures in the accompanying benchmark tiny expert sample; tasks replicated specified views rather than discovering analytical questions; strongest gains were in a specialized scientific grammar
30 visualization design-study researchers across expertise levels (Ruan et al., 2025) self-chosen LLMs across nine stages of applied visualization research participants described four recurring roles: assistant, programmer, connector, and simulator; repetitive and implementation work was the most broadly used interviews and ratings establish use and perceived need, not quality, time, or learning gains; simulated users and domain feedback require human validation

One poster sometimes cited as expertise-stratified evidence did not, in fact, recruit people at different skill levels. Ströbel and colleagues prompted GPT-4 as if requests came from beginner, intermediate, or advanced users, generated 27 visualizations, and had one expert rate them. Only three met the paper’s combined high-quality/value threshold and none met its formatting threshold. This is evidence about prompt conditions and one rater, not evidence that people at three expertise levels benefited.

Gains by capability profile

The evidence does not support a winner-take-all expertise curve. It supports a conditional map.

Capability profile Most credible present gain Why it can work Recurring cost or risk Evidence confidence
Low visualization and low implementation fluency access to a first chart, explanation, code path, and more candidate ideas natural language removes blank-page and syntax barriers cannot specify hidden choices or recognize semantic and design defects; repair instructions may remain too abstract Moderate for access; low for correct completed work
Learner with some data or tool familiarity faster mechanics, higher engagement, examples, and feedback during practice the learner can connect suggestions to a taught framework and inspect some state assisted performance and confidence can outrun independent skill; suggestions can narrow exploration Moderate for engagement and short-task speed; low for transfer
Intermediate visualization practitioner critique, alternatives, unfamiliar implementation, and representation bridging enough knowledge exists to filter advice, but the tool still removes costly search and execution work generic feedback, prompt overhead, and poor local edits Promising but based on very small cells
Visualization expert multiplication on bounded, reversible work: variants, repetitive changes, translation, debugging, and specialized implementation strong error models and direct-work baselines make delegation selective opportunity cost; loss of precise control; shallow advice; system cannot see tacit project judgment Moderate in bounded systems; low for open production practice
Domain expert with limited visualization or coding fluency translation of domain intent into a query, table, code candidate, or familiar chart the person can judge local meaning while the system supplies implementation valid-looking transformations can hide measure, join, or uncertainty errors; the person may not audit code Plausible and important; thin direct evidence
Analyst or developer with weak domain knowledge rapid profiling, code, and visual exploration strong implementation skill makes outputs inspectable patterns can be technically correct and contextually meaningless; open-ended branching can create false discovery Moderate for mechanics; low for semantic success
Multidisciplinary team faster prototypes, shared artifacts, translation between design, code, and domain language multiple people can contribute complementary checks generated artifacts can conceal ownership and handoff assumptions; production evidence is sparse Low to moderate; mostly interview and prototype evidence
Reader proactive explanation and question scaffolds can improve comprehension in some controlled tasks assistance can direct attention and connect unfamiliar form to a question a fluent explanation can amplify a misleading chart; accessibility, device, trust, and decision outcomes remain undermeasured Moderate for selected comprehension tasks; low in delivered practice

The often-proposed “intermediate sweet spot” is a reasonable hypothesis, not a settled law. Intermediate practitioners have enough knowledge to act on advice and enough remaining friction for the advice to save work. The direct evidence, however, includes only four intermediate designers in the strongest visualization-specific study.

What AI substitutes, scaffolds, amplifies, and makes newly necessary

More substitutable now

These are still consequential when generated badly, but they can often be executed and checked against visible state.

Better treated as scaffolds

A scaffold should make the person’s reasoning more visible and eventually support independent action. If it merely returns an answer, it may improve the artifact without improving the skill.

Amplified rather than replaced

AI can supply candidates and checks in all of these areas. The human capability determines whether those candidates become useful work.

Newly necessary

These requirements argue for better interfaces and defaults, not an infinite curriculum in prompt engineering. Instructions that merely teach users how to compensate for a transient model weakness should decay as models and harnesses improve. Skills that expose state, preserve alternatives, define acceptance, and support verification remain useful even as raw generation improves.

What may hurt—and what has actually been demonstrated

The evidence should separate observed losses from plausible concern.

Effect Evidence now Appropriate conclusion
Uncalibrated confidence confidence and satisfaction remained fairly high in the 20-novice study despite pervasive flaws, failed repairs, and incorrect interpretations Demonstrated in one current controlled setting. Do not use confidence as a correctness signal.
Premature acceptance experienced analysts declared completion with unresolved issues 31 times in 108 forced-error episodes Demonstrated under engineered errors. Expertise and interface structure did not eliminate acceptance risk.
Narrower exploration students in one 30-person exercise selected less diverse encodings when given AI suggestions Demonstrated narrowly. The result supports testing anchoring and diversity, not a universal creativity-loss claim.
Generic or context-poor advice practitioners preferred human experts on context, actionability, accuracy, and adaptability; professional-practice research shows why situated judgment matters Repeated descriptive evidence. Generic critique is useful as option generation, not authoritative design judgment.
Failed repair and regeneration churn the novice study counted high failure among attempts to fix clutter and unusable charts Demonstrated in a small current study. Local controls and explicit defect checks matter more than another unconstrained rewrite.
Skill atrophy or dependency participants and educators voice concern; sequential workshop studies cannot isolate it Not established. No convincing longitudinal visualization study measures delayed independent performance after sustained AI use.
Durable learning engagement, confidence, AI-assisted projects, and a small pre/post effect are reported Not established at the desired standard. Delayed unassisted transfer is largely missing.

The absence of longitudinal evidence does not make atrophy impossible. It makes it a research question rather than a conclusion.

Design implications

  1. Represent expertise as a profile. Ask about data, domain, visualization, tool, delivery, and AI fluency separately. Adapt explanation, visible state, and control to the task-relevant deficit.
  2. Preserve at least one verification path the person already understands. Show source values, intermediate tables, code, query, chart specification, or direct manipulation. Translation is useful only if the user can land in a representation they can judge.
  3. Make hidden choices explicit for people with less construction skill. Surface aggregation, filters, ordering, missing values, scale, comparison, and uncertainty before treating the chart as complete.
  4. Give experienced practitioners local control and an exit. Broad regeneration destroys deliberate work. Support precise edits, preserved alternatives, rollback, direct code or specification access, and easy return to manual work.
  5. Use critique to teach an inspection process, not only to return defects. Direct attention to evidence and ask for a judgment before revealing a recommendation. This is more likely to support learning and calibration.
  6. Separate artifact assistance from learning assistance. A production tool may legitimately optimize accepted output. An educational tool must also test independent reasoning and transfer.
  7. Do not simulate expertise when the claim concerns people. Persona prompting can test model sensitivity. It cannot substitute for recruiting people with declared capability profiles.
  8. Evaluate the reader separately. The creator’s expertise, confidence, and speed do not establish comprehension, trust calibration, accessibility, or action on the delivered artifact.

The next studies that would resolve the question

The field does not need another broad preference survey before it needs these comparisons.

1. Capability-profiled field trial

Recruit enough participants to cross visualization literacy, data or statistical skill, domain knowledge, implementation skill, and prior AI use. Use representative tasks in journalism, BI, science, and operational analysis. Compare direct work, the current general model with its normal harness, and the same system with an explicit scaffold. Measure accepted correctness, total time, repair, abandonment, and reader outcome.

2. Delayed learning and atrophy study

Randomize learners to conventional instruction, answer-oriented AI, and metacognitive AI that requires the learner to predict, inspect, and explain. Measure assisted performance, immediate unassisted performance, delayed retention, and transfer to new data and an unfamiliar chart. Log whether AI changes the diversity of explored alternatives.

3. Expertise-sensitive critique study

Give novice, intermediate, and expert creators the same correct, misleading, and ambiguous charts. Compare generic critique, context-grounded critique, deterministic checks, and human expert feedback. Measure defect detection, false alarms, successful repair, time, confidence calibration, and the advice participants reject.

4. Representation-bridge experiment

Let participants verify through prose, source table, intermediate table, SQL, Python, chart specification, or direct manipulation. Test whether routing to a familiar representation improves detection and correction, and whether it changes by data, code, or visualization expertise.

5. Production and reader follow-through

Start with real commissions and continue through authenticated delivery, mobile and assistive use, stakeholder acceptance, a later correction, and a data or dependency change. Measure the complete creator-to-artifact-to-reader episode. This is where access and first-draft speed either become durable value or disappear into cleanup and maintenance.

Bottom line

AI assistance currently widens who can start visualization work and compresses some of the mechanics for people who already know where they are going. It has not removed the need for visualization literacy; it has redistributed it. Construction syntax matters less. Framing, context, critique, verification, and reader responsibility matter more because a plausible artifact now arrives before the person has necessarily built the knowledge required to judge it.

The strongest near-term opportunity is therefore not universal delegation. It is capability-aware collaboration: identify which human resource is present, which is missing, which outcome matters, and which representation lets the person retain judgment. Then measure the gain at the accepted artifact, independent learning, or reader outcome—not at the moment the first chart appears.

Source note

This synthesis gives priority to primary papers and full-text studies. It uses the 2025 visualization-literacy review as an organizing framework, then checks foundational novice, expertise, professional-practice, education, critique, verification, and specialized-authoring studies directly. Evidence from small panels and sequential workshops is described as bounded or directional. Self-report, artifact ratings, independent learning, and reader outcomes remain separate throughout. The evidence base is predominantly English-language and Western or university-centered; the novice-definition review and the literacy review both identify representation gaps that remain relevant here.

Update log