Human skills and banked gains in AI-assisted data visualization
Status: research snapshot, evidence cut 14 August 2026. This chapter examines what a person needs to know or be able to do when making and reading data visualizations, then asks which of those capabilities AI has demonstrably improved for different people. It is a companion to The state of AI-assisted data visualization research and The practitioner and reader experience. The companion chapter Learning data visualization when AI can make the chart follows the same capabilities across the recent past, August 2026, and the next three years, then turns the evidence into a practice allocation and learning loop.
Executive summary
Visualization expertise is not one thing, and “novice” is not one population. A person may read familiar business charts fluently but be unable to program; know a scientific domain but not visualization design; or build polished charts without being able to audit a misleading transformation. Studies that label all of these people novices or experts conceal the mechanism that AI is helping.
The most useful existing framework defines visualization literacy through four families of competence: consume a visualization, construct one, critique its evidence and design, and connect it to a domain, audience, decision, or public context. Professional work also draws on data and statistical knowledge, domain semantics, implementation skill, situated design judgment, and collaboration or delivery knowledge. AI can reduce the cost of some of these resources. It does not make the whole profile uniform.
The evidence supports five conclusions.
- The best-banked gain for less experienced people is access, not autonomy. Natural language lets people attempt transformations, code, chart forms, and critique that were previously out of reach. In controlled education studies, AI use also reduced reported time, expanded idea lists, and increased confidence or engagement. These studies found much weaker evidence of durable visualization proficiency, independent transfer, or error detection.
- Intermediate and experienced practitioners appear to convert assistance into better work more reliably, but the evidence is small and bounded. In a 13-designer critique study, intermediate and expert designers used advice more effectively than novices. In a seven-expert scientific-visualization study, a restricted language and compiler reduced debugging and improved bounded replication work. These are promising mechanisms, not population estimates.
- AI substitutes for syntax and search more readily than for situated judgment. Professional designers combine framing, quality, appearance, tool, navigation, and composition judgments in ways shaped by the client, data, medium, precedent, and moment. Experienced analysts similarly describe exploration as opportunistic and domain-guided. Generic rules or a prompt cannot fully specify that work in advance.
- Verification is a human capability and a system property. People verify through representations they understand: source values, intermediate tables, SQL, Python, visual output, or direct controls. A current novice study found only three data-verification attempts across 60 sessions. A study of experienced analysts still recorded premature acceptance in 31 of 108 deliberately error-prone episodes. Expertise helps, but does not make an opaque or misleading workflow safe.
- There is one encouraging immediate learning result, but not yet convincing evidence of durable skill atrophy—or durable AI-taught construction skill. A 117-person randomized visual-comprehension study found that a proactive agent using scaffolded questions outperformed a passive agent and a data story after assistance was removed. The test was immediate and used the same visualization formats. Elsewhere, passive acceptance, reduced exploration, and dependence are visible risks, but delayed unassisted transfer remains rare. Claims that AI either teaches visualization broadly or destroys the skill still run ahead of the evidence.
The practical design rule is to measure a specific gain against a specific capability profile. Do not ask whether “novices benefit.” Ask whether, for example, a domain expert with low coding fluency can produce a correct, inspectable analysis faster; whether a design student can later critique a new chart without AI; or whether an experienced author can incorporate useful feedback without spending more time filtering it than editing directly.
A better model of visualization skill
The 2025 state-of-the-art review of visualization literacy reviewed 374 papers and proposed four overlapping competency themes. It is the best available standard spine because it includes more than chart reading and requires every competency to be operationalized in a domain, scenario, and audience.
| Competency | Literal question | Examples | What AI may change |
|---|---|---|---|
| Consumption | Can the person read and interpret the representation? | locate and compare values; read axes and legends; recognize patterns; interpret uncertainty | explain unfamiliar forms, answer questions, direct attention, translate a chart into prose or another representation |
| Construction | Can the person produce an effective representation? | select data; choose a form; map fields to marks and channels; implement, annotate, and revise | remove syntax barriers, propose encodings, write transformations and code, generate variants, apply repetitive changes |
| Critique | Can the person judge the chart and its production process? | identify misleading scales or aggregations; inspect source fidelity; question omissions; assess whether the design supports the claim | enumerate checks, run tests, inspect renders, surface common design failures—but also produce plausible rationalizations and false reassurance |
| Connection | Can the person relate the visualization to external meaning and action? | use domain knowledge; integrate a chart into a story; debate implications; make a decision; recognize who or what is missing | retrieve background, bridge creator and domain language, draft a narrative—without owning local definitions, consequences, or publication authority |
These are competency themes, not four stages and not four kinds of people. A journalist may need all four to integrate evidence into a story. An operational analyst may construct only familiar dashboards but make sophisticated connections to a live business process. A reader may need no implementation skill and still need strong critique and connection skills to judge a public claim.
The resources behind the four competencies
The four-part spine becomes more diagnostic when it is crossed with the resources that real work draws upon.
| Human resource | What it contributes | What happens when it is missing |
|---|---|---|
| Data and statistical knowledge | aggregation, comparison, missingness, uncertainty, method choice, and causal restraint | a well-drawn chart can answer the wrong quantitative question |
| Domain and semantic knowledge | field meaning, measure definitions, plausible ranges, institutional history, and decision consequences | the output can be structurally valid and locally false |
| Visual and perceptual knowledge | encoding, hierarchy, comparison, annotation, accessibility, and composition | the result can be correct but hard to read, misleading, or poorly matched to its audience |
| Tool and implementation knowledge | transformation, code, direct manipulation, interaction, responsive behavior, and production constraints | ideas remain unbuildable, uninspectable, brittle, or expensive to revise |
| Critical and epistemic judgment | verification, source assessment, alternative explanations, stopping rules, and calibrated trust | speed and polish become substitutes for evidence |
| Situated design judgment | framing, quality, appearance, tool choice, navigation, and composition in the specific project | generic best practice displaces the actual purpose, material, client, or medium |
| Collaboration, audience, and delivery knowledge | eliciting requirements, handoff, governance, accessibility, publication, maintenance, and reader testing | an acceptable draft fails as shared or delivered work |
| AI interaction literacy | scoping, specifying intent, exposing assumptions, comparing alternatives, inspecting state, correcting locally, and refusing or rolling back | the person either underuses a useful system or accepts a fluent but ungrounded result |
The first seven resources predate generative AI. The eighth is new, but should not be confused with prompt ornament. It is primarily the ability to establish an explicit work contract and evaluate what the system did.
“Novice” and “expert” are unsafe shortcuts
Burns and colleagues analyzed 79 influential visualization papers that referred to novices, non-experts, laypeople, or the general public. They found disagreement within and between papers about what defined a novice, plus a mismatch between broad claims and study samples concentrated among young people, students, and US residents. Papers variously used profession, scientific proximity, technical skill, visualization experience, and the absence of some knowledge as the boundary.
That problem is acute in AI research. A model can remove a programming barrier while leaving visualization, data, and domain barriers untouched. It can also help an experienced coder who is a novice in the dataset, or a subject-matter expert who has never built an interactive chart. A single expertise label makes these distinct gains look equivalent.
A defensible study should therefore publish a capability profile rather than a rank alone:
- visualization consumption, construction, critique, and connection;
- data and statistical experience;
- domain familiarity;
- programming and tool familiarity;
- professional practice and delivery experience;
- prior AI use; and
- the task, medium, audience, stakes, and acceptable evidence.
The label can remain as a compact description after those dimensions are visible. It cannot replace them.
What expertise looks like in observed work
Three pre-AI lines of research clarify what assistance is trying to augment.
- In How Information Visualization Novices Construct Visualizations, nine novices worked through a human mediator. Their activity centered on selecting data attributes, choosing a visual template, and specifying visual mappings. They struggled to translate questions into attributes and mappings, supplied partial specifications, and reached for familiar chart types. The study is small and from 2010, but its decomposition remains useful: natural language does not erase the choices hidden inside “make a chart.”
- An eye-tracking study across five levels of science expertise observed 36 participants completing 26 graph tasks. More expert participants allocated more attention to titles, variables, legends, and plotted data and used more directed search sequences; less experienced participants spent more attention on questions and answer choices. The study measures science-domain progression rather than visualization expertise in isolation, which is precisely the point: relevant domain knowledge changes how people see a graph.
- Interviews with 20 professional visualization designers and a focused design-judgment analysis of ten of those practitioners found little support for one orderly professional process. Practitioners used precedent and experience and layered framing, appreciative, appearance, quality, instrumental, navigational, compositional, and connective judgments in response to the situation. This is not irrationality. It is how experts move through a problem that cannot be completely specified in advance.
The implication for AI is not that guidance is useless. It is that a universal recipe can support elementary construction while conflicting with the situated, layered judgment of advanced practice.
What counts as a banked gain
The phrase banked gain is intentionally stricter than “the participant liked the system” or “the model produced a chart.” A gain is banked when it survives at the outcome level that matters for the task.
| Gain | Evidence that counts | What does not establish it |
|---|---|---|
| Access | the person successfully attempts a task or representation they could not complete under the declared baseline | a nonblank render or stated empowerment |
| Productivity | less total human and machine time or effort to an accepted artifact of comparable quality | time to first draft, self-reported ease, or omitted verification time |
| Artifact quality | a better accepted output under a declared correctness, usefulness, design, or delivery rubric | polish, one model judge, or satisfaction alone |
| Learning | better unassisted post-test, delayed retention, or transfer to a new task | confidence, engagement, AI-assisted final work, or prompt skill alone |
| Judgment and verification | more defects detected and correctly repaired, better calibration, or fewer false acceptances | more explanations or perceived control |
| Reader outcome | better comprehension, recall, decision, trust calibration, accessibility, or appropriate action | creator success or model chart-question answering |
| Sustained work | maintainability, later correction, handoff, reuse, or regression performance | one session or a published screenshot |
A system can bank one gain while losing another. Faster completion with less diverse exploration is a productivity gain and a possible learning or creativity cost. Higher confidence paired with pervasive errors is an experience gain and a calibration failure. The outcomes should not be averaged into one “AI helps” score.
Where gains have actually been observed
The table is organized by people and outcomes, not by product. “Measured” means the study observed behavior or assessed artifacts; it does not mean the result generalizes beyond that study.
| Population and study | Assistance | Gain that was observed | What was not banked |
|---|---|---|---|
| 95 masters students in a group data-story exercise (Ahmad & Ma, 2024) | general LLMs, including a 30-minute prompting lesson | the non-LLM groups reported spending 44% more time; AI users reported more focus on insights; groups with prior LLM experience received higher scores in a secondary analysis | most AI versus non-AI artifact differences were not statistically significant; non-AI groups scored slightly higher on creativity; non-use was not verified; time was self-reported |
| 30 graduate students over three sequential workshops (More Than Chatting) | peer learning, then required LLM use and prompt instruction, then optional use | confidence, perceived utility, and engagement increased; pre/post competency effect was small to moderate (d=.358); expert scores improved modestly over the sequence | no parallel control separated AI from instruction, practice, task, or workshop order; the paper itself reports a gap between engagement and skill |
| 117 higher-education participants in a randomized comprehension experiment (Yan et al., 2024) | data story, passive conversational agent, or proactive agent using scaffolded questions | all conditions improved; the proactive agent produced better immediate comprehension after assistance was removed | same visualization formats at post-test; immediate rather than delayed transfer; comprehension rather than construction |
| 26 students across four semester projects (Kim et al., 2024) | general conversational assistant used with Tableau, D3, and Vega-Lite work | reported speed, confidence, coding help, debugging, and access outside office hours; 3,773 queries document actual use | no parallel control or delayed transfer test; query themes, volume, and length were not significantly related to grades; design help was less central |
| 30 graduate students in a complementary encoding exercise (Ahmad & Ma, 2024) | LLM encoding suggestions versus students’ own knowledge | more idea and phrasing options were reported | only 14 completed the post-survey; students chose less diverse visual encodings when given AI suggestions; perceived creativity was neutral |
| 20 visualization novices across 60 sessions (Vibe Visualizing) | current general assistant used for authoring and interpretation | access, confidence, and satisfaction: participants produced artifacts and rated mean confidence 3.73/5 and satisfaction 3.93/5 | all charts had flaws; data verification appeared only three times; 11/16 clutter-removal repairs and 8/9 unusable-chart repairs failed; confidence was not calibrated to quality |
| 13 designers: six novice, four intermediate, three expert (Visualizationary) | LLM and deterministic perceptual critique, hierarchy, and version history on self-selected work | external experts rated final change 3.69/5 on average; intermediate and expert means were higher than the novice mean; 68 critiques were used selectively | no no-critique baseline; roughly 90–150 observed minutes per participant; some participants did not end on their best-rated version; no production or maintenance outcome |
| 12 practitioners bringing their own work (Kim et al., 2025) | GPT-3.5 advice compared with seasoned visualization experts | rapid brainstorming, broad option generation, and a neutral first pass | practitioners preferred human experts for accuracy, helpfulness, reliability, context, adaptability, and actionable conversation; raw model capability is dated |
| 18 experienced analysts across 108 forced-error episodes (Shang et al., 2024) | conversational, stepwise, or phasewise AI analysis interfaces | structured decomposition improved perceived control | no detected difference in task success, time, or verification hints; seven episodes were unfinished and 31 were accepted with an issue remaining |
| Seven visualization experts on three scientific replication tasks (Raiven, 2026) | restricted visualization language, deterministic compiler, and visual feedback | bounded ease and efficiency in reproducing specified scientific views; five preferred the system and two had no preference; compiler-based generation sharply reduced execution failures in the accompanying benchmark | tiny expert sample; tasks replicated specified views rather than discovering analytical questions; strongest gains were in a specialized scientific grammar |
| 30 visualization design-study researchers across expertise levels (Ruan et al., 2025) | self-chosen LLMs across nine stages of applied visualization research | participants described four recurring roles: assistant, programmer, connector, and simulator; repetitive and implementation work was the most broadly used | interviews and ratings establish use and perceived need, not quality, time, or learning gains; simulated users and domain feedback require human validation |
One poster sometimes cited as expertise-stratified evidence did not, in fact, recruit people at different skill levels. Ströbel and colleagues prompted GPT-4 as if requests came from beginner, intermediate, or advanced users, generated 27 visualizations, and had one expert rate them. Only three met the paper’s combined high-quality/value threshold and none met its formatting threshold. This is evidence about prompt conditions and one rater, not evidence that people at three expertise levels benefited.
Gains by capability profile
The evidence does not support a winner-take-all expertise curve. It supports a conditional map.
| Capability profile | Most credible present gain | Why it can work | Recurring cost or risk | Evidence confidence |
|---|---|---|---|---|
| Low visualization and low implementation fluency | access to a first chart, explanation, code path, and more candidate ideas | natural language removes blank-page and syntax barriers | cannot specify hidden choices or recognize semantic and design defects; repair instructions may remain too abstract | Moderate for access; low for correct completed work |
| Learner with some data or tool familiarity | faster mechanics, higher engagement, examples, and feedback during practice | the learner can connect suggestions to a taught framework and inspect some state | assisted performance and confidence can outrun independent skill; suggestions can narrow exploration | Moderate for engagement and short-task speed; low for transfer |
| Intermediate visualization practitioner | critique, alternatives, unfamiliar implementation, and representation bridging | enough knowledge exists to filter advice, but the tool still removes costly search and execution work | generic feedback, prompt overhead, and poor local edits | Promising but based on very small cells |
| Visualization expert | multiplication on bounded, reversible work: variants, repetitive changes, translation, debugging, and specialized implementation | strong error models and direct-work baselines make delegation selective | opportunity cost; loss of precise control; shallow advice; system cannot see tacit project judgment | Moderate in bounded systems; low for open production practice |
| Domain expert with limited visualization or coding fluency | translation of domain intent into a query, table, code candidate, or familiar chart | the person can judge local meaning while the system supplies implementation | valid-looking transformations can hide measure, join, or uncertainty errors; the person may not audit code | Plausible and important; thin direct evidence |
| Analyst or developer with weak domain knowledge | rapid profiling, code, and visual exploration | strong implementation skill makes outputs inspectable | patterns can be technically correct and contextually meaningless; open-ended branching can create false discovery | Moderate for mechanics; low for semantic success |
| Multidisciplinary team | faster prototypes, shared artifacts, translation between design, code, and domain language | multiple people can contribute complementary checks | generated artifacts can conceal ownership and handoff assumptions; production evidence is sparse | Low to moderate; mostly interview and prototype evidence |
| Reader | proactive explanation and question scaffolds can improve comprehension in some controlled tasks | assistance can direct attention and connect unfamiliar form to a question | a fluent explanation can amplify a misleading chart; accessibility, device, trust, and decision outcomes remain undermeasured | Moderate for selected comprehension tasks; low in delivered practice |
The often-proposed “intermediate sweet spot” is a reasonable hypothesis, not a settled law. Intermediate practitioners have enough knowledge to act on advice and enough remaining friction for the advice to save work. The direct evidence, however, includes only four intermediate designers in the strongest visualization-specific study.
What AI substitutes, scaffolds, amplifies, and makes newly necessary
More substitutable now
- recalling syntax and chart-library APIs;
- boilerplate transformation and rendering code;
- translating a known operation between prose, SQL, Python, and a table;
- drafting common labels, summaries, accessibility descriptions, and documentation;
- producing cheap variations or reproducing a specified view; and
- locating a conventional starting point.
These are still consequential when generated badly, but they can often be executed and checked against visible state.
Better treated as scaffolds
- decomposing a question into attributes, transformations, encodings, and checks;
- explaining an unfamiliar visualization;
- proposing alternatives and counterarguments;
- critiquing common perceptual and integrity failures;
- prompting a learner to inspect titles, variables, axes, legends, uncertainty, and source data; and
- bridging a visualization researcher and domain expert’s vocabularies.
A scaffold should make the person’s reasoning more visible and eventually support independent action. If it merely returns an answer, it may improve the artifact without improving the skill.
Amplified rather than replaced
- framing the analytical and communication problem;
- applying domain semantics and recognizing implausible results;
- choosing which evidence and uncertainty matter;
- situated design judgment and editorial voice;
- audience, accessibility, and device judgment;
- verification and calibrated stopping; and
- responsibility for delivery, disclosure, correction, and maintenance.
AI can supply candidates and checks in all of these areas. The human capability determines whether those candidates become useful work.
Newly necessary
- specifying an inspectable intent and acceptance contract;
- recognizing which choices the system made implicitly;
- comparing generated alternatives without anchoring on the first;
- choosing a familiar verification representation;
- distinguishing model explanation from evidence;
- preserving provenance, versions, and local corrections;
- knowing when deterministic tests or a specialist are needed; and
- refusing, rolling back, or switching to direct work when assistance is negative value.
These requirements argue for better interfaces and defaults, not an infinite curriculum in prompt engineering. Instructions that merely teach users how to compensate for a transient model weakness should decay as models and harnesses improve. Skills that expose state, preserve alternatives, define acceptance, and support verification remain useful even as raw generation improves.
What may hurt—and what has actually been demonstrated
The evidence should separate observed losses from plausible concern.
| Effect | Evidence now | Appropriate conclusion |
|---|---|---|
| Uncalibrated confidence | confidence and satisfaction remained fairly high in the 20-novice study despite pervasive flaws, failed repairs, and incorrect interpretations | Demonstrated in one current controlled setting. Do not use confidence as a correctness signal. |
| Premature acceptance | experienced analysts declared completion with unresolved issues 31 times in 108 forced-error episodes | Demonstrated under engineered errors. Expertise and interface structure did not eliminate acceptance risk. |
| Narrower exploration | students in one 30-person exercise selected less diverse encodings when given AI suggestions | Demonstrated narrowly. The result supports testing anchoring and diversity, not a universal creativity-loss claim. |
| Generic or context-poor advice | practitioners preferred human experts on context, actionability, accuracy, and adaptability; professional-practice research shows why situated judgment matters | Repeated descriptive evidence. Generic critique is useful as option generation, not authoritative design judgment. |
| Failed repair and regeneration churn | the novice study counted high failure among attempts to fix clutter and unusable charts | Demonstrated in a small current study. Local controls and explicit defect checks matter more than another unconstrained rewrite. |
| Skill atrophy or dependency | participants and educators voice concern; sequential workshop studies cannot isolate it | Not established. No convincing longitudinal visualization study measures delayed independent performance after sustained AI use. |
| Durable learning | engagement, confidence, AI-assisted projects, and a small pre/post effect are reported | Not established at the desired standard. Delayed unassisted transfer is largely missing. |
The absence of longitudinal evidence does not make atrophy impossible. It makes it a research question rather than a conclusion.
Design implications
- Represent expertise as a profile. Ask about data, domain, visualization, tool, delivery, and AI fluency separately. Adapt explanation, visible state, and control to the task-relevant deficit.
- Preserve at least one verification path the person already understands. Show source values, intermediate tables, code, query, chart specification, or direct manipulation. Translation is useful only if the user can land in a representation they can judge.
- Make hidden choices explicit for people with less construction skill. Surface aggregation, filters, ordering, missing values, scale, comparison, and uncertainty before treating the chart as complete.
- Give experienced practitioners local control and an exit. Broad regeneration destroys deliberate work. Support precise edits, preserved alternatives, rollback, direct code or specification access, and easy return to manual work.
- Use critique to teach an inspection process, not only to return defects. Direct attention to evidence and ask for a judgment before revealing a recommendation. This is more likely to support learning and calibration.
- Separate artifact assistance from learning assistance. A production tool may legitimately optimize accepted output. An educational tool must also test independent reasoning and transfer.
- Do not simulate expertise when the claim concerns people. Persona prompting can test model sensitivity. It cannot substitute for recruiting people with declared capability profiles.
- Evaluate the reader separately. The creator’s expertise, confidence, and speed do not establish comprehension, trust calibration, accessibility, or action on the delivered artifact.
The next studies that would resolve the question
The field does not need another broad preference survey before it needs these comparisons.
1. Capability-profiled field trial
Recruit enough participants to cross visualization literacy, data or statistical skill, domain knowledge, implementation skill, and prior AI use. Use representative tasks in journalism, BI, science, and operational analysis. Compare direct work, the current general model with its normal harness, and the same system with an explicit scaffold. Measure accepted correctness, total time, repair, abandonment, and reader outcome.
2. Delayed learning and atrophy study
Randomize learners to conventional instruction, answer-oriented AI, and metacognitive AI that requires the learner to predict, inspect, and explain. Measure assisted performance, immediate unassisted performance, delayed retention, and transfer to new data and an unfamiliar chart. Log whether AI changes the diversity of explored alternatives.
3. Expertise-sensitive critique study
Give novice, intermediate, and expert creators the same correct, misleading, and ambiguous charts. Compare generic critique, context-grounded critique, deterministic checks, and human expert feedback. Measure defect detection, false alarms, successful repair, time, confidence calibration, and the advice participants reject.
4. Representation-bridge experiment
Let participants verify through prose, source table, intermediate table, SQL, Python, chart specification, or direct manipulation. Test whether routing to a familiar representation improves detection and correction, and whether it changes by data, code, or visualization expertise.
5. Production and reader follow-through
Start with real commissions and continue through authenticated delivery, mobile and assistive use, stakeholder acceptance, a later correction, and a data or dependency change. Measure the complete creator-to-artifact-to-reader episode. This is where access and first-draft speed either become durable value or disappear into cleanup and maintenance.
Bottom line
AI assistance currently widens who can start visualization work and compresses some of the mechanics for people who already know where they are going. It has not removed the need for visualization literacy; it has redistributed it. Construction syntax matters less. Framing, context, critique, verification, and reader responsibility matter more because a plausible artifact now arrives before the person has necessarily built the knowledge required to judge it.
The strongest near-term opportunity is therefore not universal delegation. It is capability-aware collaboration: identify which human resource is present, which is missing, which outcome matters, and which representation lets the person retain judgment. Then measure the gain at the accepted artifact, independent learning, or reader outcome—not at the moment the first chart appears.
Source note
This synthesis gives priority to primary papers and full-text studies. It uses the 2025 visualization-literacy review as an organizing framework, then checks foundational novice, expertise, professional-practice, education, critique, verification, and specialized-authoring studies directly. Evidence from small panels and sequential workshops is described as bounded or directional. Self-report, artifact ratings, independent learning, and reader outcomes remain separate throughout. The evidence base is predominantly English-language and Western or university-centered; the novice-definition review and the literacy review both identify representation gaps that remain relevant here.
Update log
- 2026-08-14 — Initial public edition. Defined the human capability profile, separated seven kinds of banked gain, and assessed which gains are observed for novices, intermediate practitioners, and experts.