Research snapshot · August 2026

The practitioner and reader experience of AI-assisted data visualization.

AI is already good at getting someone from a blank page to a plausible first thing. It is much less reliable at carrying the work through local meaning, correction, delivery, and audience understanding.

Short answer

The reliable win is compression, not delegation. Faster creation matters; it does not transfer responsibility for the analytical claim or the reader’s understanding.

For creators, readers, and product teamsThis companion translates research results into the lived jobs, gains, costs, and trust questions surrounding current tools.

Not a product rankingProducts appear to explain interaction environments. Documentation establishes features, not comparative quality.

The central gap

A generated chart can be possible without becoming finished work—or useful communication.

Most demos answer the first question below. Practitioners live in the second. Readers determine the third.

Who is reaching for what

AI is entering through existing work, not replacing the visualization stack.

The best available adoption evidence shows practitioners adding AI beside spreadsheets, code, design, and BI tools. Creation and analysis are ahead of trusted conversational consumption.

Data Visualization Society · 2024

825 practitioners answered whether they had used AI in visualization work during the prior year.

Online, self-selected survey. The result describes respondents, not the population or product market share.

Use was up 13 percentage points from the 2023 survey. The 28 unsure responses are the narrow violet segment.

Among 305 AI users

Early-lifecycle jobs were more common than visualization production.

Multiple selections were allowed, so the bars do not sum to 100%.

Prepare or clean184
Analyze134
Ideate or storyboard119
Produce visualizations92
Support a viz team27
WhoFirst reachJobEvidence

Visualization practitioners

A general assistant beside Excel, Tableau, Python, R, D3, Figma, or a publishing tool

Prepare, debug, learn, storyboard, draft labels, make a first view

Observed survey patternSpecific AI-product share was not measured.

Analytics engineers and code-capable analysts

General coding assistance, then project-aware notebook or development agents

Write SQL or Python, document models, debug pipelines, create analysis output

Observed survey patternBroader than visualization and vendor-adjacent.

Spreadsheet-native teams

AI inside Excel or Sheets; AI spreadsheets when code or live connections outgrow the grid

Ask about a range, create formulas and charts, preserve a familiar handoff

Installed-base inferenceExcel use is observed; AI-feature adoption is not.

BI authors and data teams

AI inside an existing BI, notebook, or governed data platform

Author reports, reuse measures, inspect queries, answer follow-ups, govern access

Provider target + casesNo independent head-to-head field test.

Business consumers and executives

Conversational BI over curated metrics

Retrieve a number, ask why it changed, get an ad hoc cut without navigating a report

Provider target + testimonyTrust depends on narrow, owned definitions.

Designers, journalists, and communicators

General help for ideation or code, then a deliberate design and publishing surface

Explore, create variants, annotate, improve accessibility, implement a story

Small subgroup + inferenceNot a population estimate.

DVS survey and public data ↗ · Analytics-engineering survey ↗. Both are self-reported and nonrepresentative; denominators remain attached.

The jobs

Visualization is a workflow, not a prompt.

“Make a chart” hides work before, during, and after visual encoding. The same assistant can be strong at one job and harmful at the next.

01

Define

Decide what is worth doing and which data is allowed to answer it.

  1. Frame the questionRestate, clarify, propose hypothesesHuman owns the decision and the refusal.
  2. Acquire and governFind tables, describe schemas, draft access stepsHuman owns authority, permission, and allowed use.
02

Work the data

Turn source material into evidence without losing its meaning.

  1. Prepare and transformClean, join, reshape, calculateHuman owns definitions and exclusions.
  2. AnalyzeSummarize, compare, model, find anomaliesHuman owns method and uncertainty.
  3. ExploreBranch into cheap views and follow-upsHuman owns relevance and stopping.
03

Make the artifact

Choose what the audience will see and make every consequential choice inspectable.

  1. Sketch or reproduceMake a candidate, mockup, or reference copyHuman owns whether the reference fits.
  2. Choose and encodeSelect aggregation, form, emphasis, annotationHuman owns the analytical and rhetorical choice.
  3. Critique and refineImprove layout, labels, interaction, accessHuman resolves intent and conflicting advice.
  4. Verify and debugInspect data, code, state, values, and claimsHuman defines decisive acceptance checks.
04

Deliver and sustain

Test the real surface, support the reader, and keep the work correct later.

  1. Package and publishAssemble, document, export, deployHuman owns privacy and release gates.
  2. Explain and interrogateAdd summaries, tooltips, guided questionsHuman owns the source claim and reader test.
  3. Maintain and updateRegenerate, detect drift, hand off, retireHuman owns semantic and dependency changes.

Listening to the field

People are not having one AI visualization experience.

The same person can be delighted by a first draft, frustrated by the repair loop, afraid to put their name on the result, and still choose to use the tool tomorrow. Attitude and behavior do not move together.

01 · Before the work

Is this invitation, pressure, or a threat?

People arrive with organizational demands, prior skill, career hopes, and fears about what assistance may remove.

Analytics team lead · workplace mandate

“We need to use AI” arrived before a useful job.

Leadership wanted visible thought leadership. The custom chatbot was unreliable; for the immediate task, a pivot table was faster. The frustration was having to perform adoption.

Visualization coder · personal project

Fear of losing the rewarding part gave way to momentum.

Jisell Howe first worried that instant code would remove the journey from idea to customized chart. A concrete build changed the feeling: errors remained, but search and troubleshooting became less disruptive.

Master’s student · guided practical

Early anxiety became a sense of access and accomplishment.

AI produced a first bar chart, but an instructor supplied the visualization knowledge needed to improve it. The gain was reaching the work—not yet independent mastery.

02 · First candidate

The initial rush is real—and often comes from staying in motion.

Delight tends to attach to access, flow, and continuity across chores, not only to a chart appearing.

Independent creator · personal dashboard

Loose instructions worked because the materials were ready.

A two-day revision felt surprisingly fluid. Four short lines were enough once the current files, requested changes, and source URLs were present. The creator’s lesson was that organized context mattered more than prompt polish.

Technologist · visualization class project

The pleasure came from continuity across a whole project.

Sef Kloninger called the work “plain fun.” The agent moved with him from a 500 GB dataset through tests, exploration, a dashboard, debugging, and a presentation; he still inspected raw data and requested checks.

Experienced BI analyst · unfamiliar API

A week-sized unknown became a day-sized job.

The practitioner immediately added the caveat: prior proficiency made the compression possible. A novice in another account valued a smaller gain—producing cleaning and validation scripts that had been out of reach.

03 · Working loop

Assistance turns into interruption when precise repair begins.

The experience depends on whether explaining, waiting, inspecting, and retrying costs less than direct manipulation.

Power BI practitioners · report authoring

Greenfield structure was fast; local refinement was clumsy.

Stale filters, bookmark identifiers, whitespace churn, and slow edit-preview cycles accumulated. One experienced practitioner said describing a small change could take longer than making it.

Analytics educator and practitioner · stakeholder work

“Babysitting” meant protecting a professional reputation.

Christina Stathopoulos described useful brainstorming and exploration, then recalled a generated chart whose accompanying interpretation reversed the visible comparison. The feared failure was looking careless in front of a stakeholder.

Data journalist · published interactive

The build became hardest when it looked almost finished.

A recognizable site appeared in about an hour. Geography, evidence notes, rate limits, mobile behavior, architecture, and fact-checking then consumed repeated rounds of repair.

04 · Acceptance

Trust becomes personal when someone must accept the work.

The imagined stakeholder, patient, client, colleague, or excluded reader changes which errors matter and where refusal is rational.

Dashboard consultant · client reconstruction

A 25-minute reconstruction became an existential pricing question.

An agent recreated roughly a week’s visible dashboard work from a finished screenshot and source tables. The consultant’s question was what clients had really paid for: construction, diagnosis, or the judgment embedded in the reference.

17 biomedical-visualization practitioners · consequential work

Refusal was sometimes expertise, not resistance to change.

Participants welcomed boilerplate, inspiration, and translation while some rejected final AI imagery for anatomy, patient communication, or scientific work. Others protected rendering because it was also where control, flow, and creative joy lived.

Blind journalist · news reader and colleague

Access felt different when it became a shared habit.

Johny Cassidy describes exclusion from charts with missing or useless descriptions—and the relief of colleagues taking responsibility. Generated alt text is not accessible delivery without alternatives, feedback, and correction.

05 · After delivery

Creation stories are vivid. The artifact’s afterlife is mostly quiet.

Use, maintenance, correction, learning, and second-person handoff determine whether the saved effort became value.

BI practitioner · urgent dashboard request

Praise and distribution did not become use.

A manager praised an urgent dashboard and sent it to six colleagues. Three months later, the usage record showed one manager view. Creator pride, stakeholder approval, and reader use were three different outcomes.

BI developer · adjacent work

The durable gain sat around the chart.

AI documented SQL and calculations, reviewed junior work, and rehearsed stakeholder questions. Those tasks made the dashboard easier to explain and maintain without asking the model to own the final claim.

Freelance data journalist · analog practice

Some friction created attention rather than waste.

Emilia Ruzicka collected and drew personal data by hand. Slowness, imperfection, and direct contact with the data produced experimentation and care—the kind of learning a faster route can accidentally remove.

Listening past the original creator

Non-use is often quiet. Reliance is often provisional.

A second pass pursued spreadsheet-native work, people who do not make AI their default, recipients deciding whether an answer is defensible, blind and low-vision learners, and the second person asked to maintain the result.

Routine public-sector work · mixed-methods evaluation

People kept the assistant and routed experienced Excel work around it.

DWP staff allocated tasks according to expertise, time, trust, data sensitivity, and habit. Data-heavy Excel work, chart generation, and intricate formatting remained weak points. One person valued having the assistant but was often too busy to remember to use it.

DWP evaluation ↗1,716 user responses · 2,535 comparison responses · 19 interviews · nonrandom licences and no pre-trial baseline
Decision recipients · within-person comparison

The preferred interface changed with the acceptance criterion.

When speed mattered

Chat 15Dashboard 3Both 2

When confidence mattered

Chat 0Dashboard 18Both 2

Twenty-participant exploratory study ↗. The chatbot compressed retrieval; the dashboard retained overview and an inspectable chain to the data. Eighteen participants had computer-science backgrounds and were proxies for industrial decision makers.
Spreadsheet-help community · public discussion

An instant private answer can also remove a public learning episode.

A supply-chain analyst noticed that AI had displaced visits to a peer forum. Replies described both useful help and invented functions, wrong references, damaged formulas, and refusal. Several people valued the incidental learning produced by solving someone else’s problem in public.

Program recipients and an executive-dashboard observer

“Go check the raw data” is not a recipient verification contract.

Program managers and funders wanted overview, drill-down, definitions, neutral language, and review support. In a separate sales-dashboard demonstration, an observer rejected the idea that an executive seeking a quick answer should independently validate regenerated charts against raw data.

Blind and low-vision chart learners · 12-participant study

A useful answer can still fail to supply the spatial model.

Eleven participants preferred tactile charts plus text and chat; one said the better mode depended on complexity; none preferred text and chat alone. Touch supplied spatial structure and chat supplied flexible clarification. Measured chart-understanding accuracy did not improve.

Read the study ↗12 participants · 263 substantive queries · short English-speaking US study, not a delivered-artifact or long-term learning evaluation
Second maintenance · public discussion

The refresh job lived on the creator’s laptop.

An AI-built dashboard failed while its creator was away. The inheriting maintainer replaced the creator-local scheduled job with a proper pipeline; replies surfaced missing metric semantics, hidden dependencies, documentation gaps, and service expectations nobody had owned.

What listening adds

Capability scores cannot tell us what the gain costs—or what people are trying to protect.

  • Delight is often momentum.Staying in motion across formerly blocking chores can matter more than a perfect first chart.
  • Frustration is often interruption.The comparison is prompt-and-repair time versus direct work, not AI versus a blank page.
  • Fear has distinct objects.Employment, reputation, learning, craft, access, scientific harm, and vendor dependence require different responses.
  • Trust has a face.People picture who will spot the mistake, bear its consequence, or be unable to inspect the claim.
  • Refusal can be expert practice.Keeping some work manual can preserve accountability, knowledge, control, or the purpose of doing it.
  • The afterlife is underreported, not empty.One direct handoff failure exposes hidden execution and ownership assumptions; routine use, cross-release maintenance, and retirement remain faint.

Field perspectives

Implementation is getting cheaper. The argument is what the saved effort should buy.

These sources expose purposes and production constraints that benchmarks omit. Most are first-person cases, interviews, or essays: they establish that a perspective or experience exists, not how common or effective it is.

TensionOne pressureCounterweightContext decides

Construction speed versus verification debt

Two current newsroom diaries describe work that once took weeks appearing in days or hours.

They also contain wrong totals, broken geography, missed notes, repeated correction, usage limits, mobile checks, and new validation work. Benn Stancil explains why a plausible chart cannot validate its calculation.

Semantic stakes, independence of the check, cost of error, and ownership of the complete pipeline.

Lower implementation barriers versus more valuable judgment

Enrico Bertini maps opportunities across data acquisition, wrangling, analysis, tours, and creative exploration. Alberto Cairo treats easier code as capacity extension.

Question quality, interpretation, and final analytical choice become more important. Exploratory, explanatory, essayistic, and artistic work do not share one quality function.

Purpose, audience, consequence, capability profile, and whether the claim remains inspectable.

Frictionless production versus productive friction

Removing syntax and formatting work can get a person to a candidate before the question goes cold.

A classroom exercise used failed prompting to reveal missing chart structure. An analog practice made slowness a source of attention, experimentation, and care.

Whether the effort is obsolete mechanics or practice in structure, skepticism, memory, and meaning.

Automated access versus independent verification

Generated descriptions and alternatives may help newsrooms address access at a scale that manual workflows have not reached.

Elavsky and Xiong Bearfield show model chains losing data, source, uncertainty, design, and purpose. Blind journalist Johny Cassidy describes an organizational and multimodal problem, not an alt-text checkbox.

Whether the reader can test the account, choose another modality, report failure, and obtain remediation.

Static charts versus conversational or divergent representations

Richard Brath asks which charts remain useful when models extract insights. Elijah Meeks proposes audience, intent, interpretation, and conversation as framework objects.

Domestic Data Streamers argues that visualization remains a distinct language for comparison, human experience, and creative divergence.

Reader task, need for overview or shared evidence, qualitative meaning, novelty, and an auditable interaction trace.

Production cases

Two current build diaries show where the time moved.

Both are self-reports by the people who built and checked their own projects. Keeping the gain beside the incurred work prevents “built in a week” from becoming a quality claim.

Lifecycle stageOttaviani: reconstructed projectsTordecilla: health dashboardNeither establishes

Baseline and first candidate

One map fell from roughly three weeks in 2012 to a two-day reconstruction.

A recognizable site appeared in about an hour; the full dashboard and checking workflow took one week.

An equal-budget comparison with the best current direct workflow or another practitioner.

Defects and correction

Incorrect totals and incomplete bilingual text required source comparison and repeated correction.

Cities rendered in the sea or broken shapes; exact evidence and source notes were sometimes missed.

A field distribution of error, correction time, regression, premature acceptance, or abandonment.

Architecture and control

About 20 ordered tasks, source checks, and open code made components easier to trace and revisit.

Extraction, data, presentation, insights, and fact-checking were separated so repairs could be replayed.

That decomposition or a model-built validator independently improves correctness across projects.

Delivery and readers

Working prototypes were published; attention, understanding, and civic use remained the stated bottleneck.

Mobile rendering was checked and stakeholders reacted; readers were not tested.

Audience understanding, accessibility, consequential use, or sustained adoption.

Sustainment

Record reconciliation and scheduled updates were proposed next steps.

Refactoring and longer-term maintenance remained after the one-week build.

Later refreshes, dependency changes, second-person handoff, incident response, or retirement.

Ottaviani build diary ↗ · Tordecilla build diary ↗. These cases support process claims, not general productivity or accuracy estimates.

The tools

The work surface determines what AI can see, change, and leave behind.

This is a showcase of current approaches, not a ranking. Product pages establish feature and target-user contracts; measured systems are labeled separately.

Feature contract

Provider or maintainer says the capability exists.

Measured system

A defined human study exists, though often on an older system version.

Conversation and artifacts

Bring a bounded file to the model—or ask it to make a small application.

Fastest route to a one-off result. Local semantics, delivery, and maintenance arrive mostly through the prompt and the person.

  • ChatGPT data analysis

    Executed analysis, tables, common charts, code, and downloads.

    Control: inspect code and intermediate results.
  • Claude Artifacts

    Generated HTML or application code for custom interactive explanation.

    Watch: state, accessibility, hosting, maintenance.
  • Julius

    Specialist data chat for files, connections, charts, models, and reports.

    Independent correctness and repair evidence is sparse.
Spreadsheet-native

Keep the familiar grid while adding analysis, code, and chart generation.

Installed context and handoff improve. Refresh behavior and copied data can quietly become the new failure surface.

  • Copilot in Excel

    Python-backed answers and optional static chart or table insertion.

    Advanced mode can create editable, refreshable Python cells.
  • Gemini in Sheets

    Summaries, formulas, prompted edits, and editable inserted charts.

    Inserted charts follow copied support data, not later changes to the original range.
  • Bricks

    One workspace for spreadsheet data, live dashboards, slides, and team collaboration.

    Evidence: provider contract.
  • Quadratic

    AI spreadsheet with Python, SQL, JavaScript, live connections, and visible code cells.

    Control: schema and code remain in the grid.
Notebooks and visible canvases

Place generated work inside an inspectable analysis state.

These environments give AI more project context while preserving cells, nodes, versions, and role-specific controls.

  • Hex AI

    Edits SQL, Python, chart, pivot, and Markdown cells; builds apps; queries curated data.

    Technical authors audit code; consumers use published apps; managers govern context.
  • Observable Canvases AI

    Adds tables, transformations, SQL, and charts beside existing work.

    Control: creates new versions instead of silently overwriting or deleting.
Guided storytelling editor

Use natural language as a shortcut into a known visual grammar.

The assistant can make precise native edits because the editor constrains the available state.

  • Flourish AI

    Changes style, labels, sources, annotation, and accessibility settings in the editor.

    Control: reversible native edits. Boundary: it does not edit source data or invent intent.
Established BI authoring

Ask AI to work inside a report, worksheet, or maintained semantic model.

More local context can narrow errors, but it also makes data teams responsible for metadata, measures, instructions, and review.

  • Power BI Copilot

    Creates and modifies report pages against model and report state.

    Depends on measures, descriptions, permissions, and review.
  • Tableau Agent

    Builds views and calculations in existing web-authoring surfaces.

    Control returns through the worksheet and data source.
  • Looker Conversational Analytics

    Maps questions onto governed LookML fields and measures.

    Administrators own glossaries, defaults, and verified queries.
Governed conversational analytics

Let business consumers ask questions over curated organizational data.

Provider direction is converging on a separate author who owns semantics, permissions, review, and feedback.

  • ThoughtSpot Spotter

    Search and conversational analysis for business teams.

    Data teams maintain governed models rather than answer queues.
  • Ask Sigma

    Shows sources, formulas, filters, and multi-step analysis.

    Control: edit individual steps instead of regenerating everything.
  • Qlik Answers

    Structured and unstructured RAG plus chart and dashboard agents.

    Includes permissions and audit logs.
  • Databricks Genie

    SQL-backed answers and visualizations in a curated no-code chat.

    Authors monitor, review, and refine definitions and instructions.
  • Amazon Q in QuickSight

    BI authoring, Q&A, executive summaries, and data stories.

    Inherits datasets, topics, permissions, and author review.
  • Oracle Analytics AI Assistant

    Charts and narratives for consumers over author-prepared data and metadata.

    Explicitly separates author and consumer roles.
Mixed-initiative and open source

Constrain generation with visible data state, declarative specifications, and reusable guardrails.

These approaches make model output more inspectable and executable. Current repositories often exceed the versions evaluated in research.

  • Data Formulator

    Direct encoding plus natural-language transformation, branching, visible derived data, and reports.

    Measured system Small reproduction study; current project is newer.
  • Lumen

    Serializable declarative pipelines, charts, dashboards, and specialist agents.

    Generated work can move from chat into notebooks or apps.
  • Vizro

    Low-code Python dashboard specification with code escape hatches and agent tools.

    Production components narrow generation; acceptance tests remain necessary.
  • AntV visualization skills

    Retrievable library syntax, chart vocabulary, declarative output, and version guardrails.

    Reports its own 174-case tests; independent current-model ablation is still needed.

Creator experience

The same capability can feel like magic at the beginning and friction at the end.

Public accounts and measured studies converge on a division of labor: coarse, repetitive, and reversible work benefits first; semantic decisions and precise refinement remain costly.

What feels amazing What appears later Durable division of labor

A first draft appears immediately. The idea becomes concrete enough to inspect and discuss.

The draft may encode an implicit aggregation, inherit a bad reference, or answer a nearby question.

AI: candidates and execution. Human: question, meaning, and selection.

Repetitive work collapses. SQL, calculations, docs, labels, bulk changes, and boilerplate move quickly.

Generated internals can be opaque; one strange implementation raises maintenance and explanation cost.

AI: legible repetition. Human: expected behavior and verification.

Many directions become cheap. A person can request alternatives without mastering every tool.

Filtering weak ideas requires expertise; novices may mistake breadth or polish for judgment.

AI: enumeration. Human: relevance, feasibility, and restraint.

A prototype improves the conversation. Stakeholders can react to layout and content before production data exists.

Prototype satisfaction says nothing about the correctness, refresh, permissions, or maintenance of the final system.

AI: negotiable mockup. Human: production contract and acceptance.

An unfamiliar representation becomes accessible. The model translates among prose, SQL, Python, tables, and charts.

A person without one familiar inspection surface may have no reliable way to evaluate the translation.

AI: translation. Human: check in a representation they understand.

Measured evidence

Speed, completion, confidence, and correctness can move in different directions.

The studies below answer different questions. Their denominators and limits remain attached; the numbers should not be pooled into one score.

Randomized public-health exercise · 2025 · 30 analyzed participants

The integrated tool was faster. Its work was less often free of serious errors.

The exercise compared integrated ChatGPT analysis with an R/Stata-plus-ChatGPT workflow on simulated epidemiological data. Overall scores were not significantly different.

Median completion timelower is better
Integrated AI38 min
Distributed tools45 min
Submissions free of serious errorshigher is better
Integrated AI6.7%
Distributed tools26%
Source ↗ Small, underpowered, 45-minute study with simulated data and no non-AI control. The bar scales are separate and labeled.
Vibe Visualizing · 2026 preprint

Novices could make charts. They could not reliably tell when the work had failed.

Twenty visualization novices completed 60 ChatGPT sessions producing 175 charts. The study observed prompting, chart quality, interpretation, verification, and repair—not just whether an image appeared.

Fatal task noncompliance52 of 60 sessions

The output omitted or contradicted a required element.

Incorrect insight recorded12 of 60 sessions

Poor chart design contributed to most incorrect-insight cases discussed.

Unusable chart22 of 175 charts

Every chart had at least one coded design flaw.

Meanwhile mean confidence was 3.73/5 and satisfaction 3.93/5. Only three verification attempts were observed.

Newer models produced fewer design flaws on replayed initial prompts, but richer interfaces added latency, broken controls, blank renders, and unverifiable interpretations. Better generation changed the failure surface; it did not remove the verification problem.

Source ↗ Small controlled datasets; the Gemini and Claude comparison replayed prompts and was not a live user study.
StudyWhat it measuredWhat it foundWhat it cannot establish

Analyst verificationCHI 2024 · 22 analysts · 52 workflows

Use of explanation, code, original and intermediate data, results, and visual summaries

People began with procedure, then often moved to data after noticing trouble; expertise shaped the artifact they trusted

Prepared tasks at one company; data transformations rather than complete visualization projects

Data Formulator 2CHI 2025 · eight participants · 16 charts each

Mixed direct manipulation and natural language on reproduction tasks

All completed; visible transformed data, code, explanations, history, and branches supported different verification styles

Six needed hints; no baseline, open exploration, self-owned data, or long-term use

Dashboard prototyping2025 preprint · 10 formative + 28 evaluation participants

Rapid mockup generation, structured edits, and comparison with a lightly taught Tableau condition

Participants valued speed, simulated data, history, and a concrete object for negotiation

Pre-data prototypes, not analytical correctness, deployment, or ongoing dashboard use

Visualization adviceTOCHI 2025 · 119 forum questions + 12 practitioners

AI versus human-expert design advice, ratings, interviews, and practitioner preference

AI helped enumerate ideas; practitioners preferred experts for accuracy, context, adaptability, and actionable advice

Practitioner sessions used 2023-era GPT-3.5; raw capability comparison is dated

Visualizationary2024 preprint · 13 designers + three expert raters

At least five versions of a self-selected visualization over a three-to-five-day window with LLM and perceptual critique

Final work improved 3.69/5 on average; intermediate and expert designers converted feedback into useful edits more readily

Roughly 90–150 minutes of observed work each; no critique baseline, production delivery, or later maintenance

Steering and verificationUIST 2024 · 18 analysts · 108 task episodes

Correction, premature acceptance, non-completion, task time, hints, and perceived control across conversational and decomposed interfaces

Seven episodes were not completed and 31 were declared complete with an issue remaining; structure improved perceived control, not detected success or time

Tasks were engineered to contain model errors and stopped at 15 minutes; no publication, maintenance, or field abandonment

Analyst verification ↗ · Data Formulator 2 ↗ · DashChat ↗ · Visualization advice ↗ · Visualizationary ↗ · Steering and verification ↗

Human capability

“Novice” and “expert” hide the skill AI is actually changing.

A person can read a familiar dashboard but not code, know the domain but not visual design, or implement polished charts without being able to audit a misleading transformation. Evaluate the relevant capability, not one rank.

01

Consume

Read values, encodings, patterns, uncertainty, and unfamiliar forms.

02

Construct

Select data, choose a form, map fields, implement, annotate, and revise.

03

Critique

Test source fidelity, hidden transformations, misleading design, and omissions.

04

Connect

Relate the chart to domain meaning, audience, story, decision, and consequence.

Every competency is contextual. Data and statistical knowledge, domain semantics, visual design, implementation, situated judgment, and delivery experience are separate resources. AI may remove one barrier while leaving the others intact.

Growing the skill

The bottleneck moved from making the chart toward judging what was made.

That does not make direct work obsolete. It changes which difficulty deserves practice. The right test is what the person can explain, inspect, repair, and transfer after assistance is removed.

Recent past

Implementation consumed the entry budget.

Tool access, syntax, debugging, scattered examples, and blank-page uncertainty kept many people from attempting the work.

Both mechanics and judgment required direct practice.
August 2026 · observed

Explanation and production are cheaper than independent judgment.

Learners report faster coding and debugging. Proactive question-based scaffolding can improve immediate post-removal comprehension. Delayed construction transfer is mostly unmeasured.

Assisted performance and learning must be scored separately.
Likely next · forecast

Direct work becomes an audit and recovery capability.

More routine implementation will be delegated. Mental models, critique, verification, local repair, and reader responsibility remain the scarce work.

Revisit if answer-oriented assistance demonstrates delayed transfer to unfamiliar tasks.
Invest more

Judgment that makes a plausible chart trustworthy.

Framing · data semantics and statistics · critique · verification and calibration · alternative comparison · domain and audience judgment · accessibility · provenance and delivery

Maintain

Material fluency for inspection and recovery.

Direct construction · data wrangling · code and specification reading · sketching · hand-checking values · debugging transformations · precise local repair

De-emphasize

Recall work whose value decays with tools and models.

API trivia · boilerplate · exhaustive taxonomy recall · manual pixel polishing · prompt incantations · deep recall of one tool's transient interface · first-render speed as a badge of skill

What “the hard way” should preserve: predict before revealing, translate questions into fields and encodings, check sample values, generate alternatives before seeing suggestions, diagnose before repair, explain decisions, and periodically transfer without assistance. Boilerplate and API hunting do not become educational merely because they are slow.

Semester-long visualization course study ↗ · Proactive scaffolding experiment ↗ · One-year visualization retention ↗ · Guardrails and unassisted learning ↗

Capability profileMost credible gainWhat remains unbankedEvidence

Low visualization + implementation fluencycurrent novice evidence

Access to a first chart, code path, explanation, and more candidate ideas

Correctness, hidden-choice detection, verification, and reliable repair

Access: moderateAccepted work: low

Learner with some data or tool fluencycourse studies plus one randomized comprehension test

Reported speed, engagement, confidence, mechanics reduction, and immediate post-removal comprehension from proactive scaffolding

Delayed independent construction and transfer to unfamiliar tasks; creativity and artifact-quality findings remain mixed or modest

Near transfer: promisingDelayed transfer: unknown

Intermediate practitionerfour-person cell in the best direct study

Turning critique, alternatives, and unfamiliar implementation into useful edits

A general “sweet spot”; the direct expertise-stratified sample is too small

PromisingNot settled

Visualization expertcritique and scientific replication

Bounded multiplication: option filtering, representation bridging, debugging, and constrained implementation

Open-ended judgment, production delivery, and a universal advantage over direct work

Bounded: moderateField: low

Domain expert, weak chart or code fluencyimportant but under-studied profile

Translation of domain intent into a query, table, code candidate, or familiar chart

Audit of joins, measures, uncertainty, interaction, and generated implementation

Direct evidence: thin

A gain is banked only at the outcome that matters.

Access is successful new work. Productivity is less total effort to an accepted artifact. Quality uses a declared correctness or usefulness rubric. Learning survives an unassisted transfer test. Verification detects and repairs defects. Reader outcome changes comprehension or decisions. Satisfaction and first-render speed do not stand in for the other rows.

The design requirement is not to pick one “user level.” Expose consequential choices for people with less construction skill, preserve precise control for experienced practitioners, and let every person verify through a representation they understand. One current experiment demonstrates immediate post-removal comprehension from proactive scaffolding; durable construction gain and AI-caused atrophy remain unestablished.

Visualization literacy review ↗ · Who counts as a novice? ↗ · Visualization learning ↗ · Creativity and time ↗ · Current novice study ↗ · Expert replication ↗

Reader experience

The reader sees a claim—not the prompts, repairs, or uncertainty behind it.

Creator delight is not a proxy for reader comprehension. The visual artifact inherits expectations from journalism, science, business, education, or social media.

Perceived trust

Clarity and familiarity matter—but neither proves truth.

In a 2025 study, 31 of 37 participants mentioned clarity while ranking charts. Sources, integrity, familiar forms, and aesthetics also mattered, with substantial differences between people. The study measured deliberative trust, not correctness, comprehension, or behavior.

Trustworthy by Design ↗
Provenance labels

“AI-generated” is not a universal trust switch.

In a preregistered 2021 experiment, people brought strong preferences for human or algorithmic recommendations, but data relevance usually dominated actual selections. The same label suggested precision to some and missing human judgment to others.

Vis Ex Machina ↗
Model evaluation

A model that decodes a chart is not a human-reader test.

On 60 synthetic charts, multimodal models enumerated data structure while 24 people formed trend narratives and reacted more to layout and overlap. Each could succeed on a different notion of intent.

How Do LLMs See Charts? ↗
Reader scaffolding

Guiding attention can outperform merely answering.

In a randomized 117-person experiment, data stories, passive AI Q&A, and proactive scaffolded dialogue all improved visual-comprehension scores. After support was removed, the proactive group’s median was 6/6 versus 5/6 for both alternatives; completion time did not differ.

GenAI agents and visual comprehension ↗
Measured harm

A data-consistent chart can still induce wrong answers.

In a 48-person controlled experiment, second-phase accuracy was 71.9% for readers shown AI-generated misleading charts and 88.3% for readers shown correct charts. This measures bounded chart-QA harm—not prevalence, calibrated trust, or consequential decisions.

ChartAttack ↗
Accessible delivery

The device, modality, chart type, and interaction all change the outcome.

A 26-person low-vision smartphone study found 100% completion with a full interactive treatment versus 61.5% with a screen magnifier. A separate 10-person blind-reader study found similar device-level accuracy but different time, workload, and chart-type results. Neither tested AI-generated charts.

Context

There is a common evidence spine, but no single correct mode of assistance.

Preserve the question, data, decisions, artifact, provenance, and acceptance evidence. Then match the authoring and evaluation method to the human purpose.

Journalism

Explain a bounded account of the evidence.

Helpful: exploration, annotation alternatives, accessible implementation, source checks.

Danger: fluent generation substituting for reporting, authorial judgment, or reader testing.

Operational dashboard

Maintain awareness and support response.

Helpful: governed measures, stable layout edits, anomaly explanation tied to source visuals.

Danger: silent changes to metrics, thresholds, state, or hierarchy.

Analytical BI

Compare evidence and make a decision.

Helpful: semantic grounding, visible intermediate data, reusable verification.

Danger: confident answers over weak metadata, ambiguous measures, or hidden filters.

Open exploration

Form and test provisional hypotheses.

Helpful: many cheap views, branches, undo, direct manipulation.

Danger: turning the first plausible pattern into the final narrative.

Science

Support specialized and reproducible claims.

Helpful: restricted representations, executable transformations, linked views, domain checks.

Danger: visual plausibility standing in for scientific correctness.

Education

Help a learner build intuition.

Helpful: adjustable assumptions, guided interaction, feedback, multiple representations.

Danger: generating interactivity without measuring what learners understand.

The lifecycle test

The same sequence exposes different evidence holes.

Question → data authority → generation → checking → correction → delivery → reader use → sustainment. The spine stays fixed; acceptance changes with the environment.

EnvironmentEvidence nowLifecycle disappearsContextual acceptance

Journalism and public explanation

Current diaries expose source work, decomposition, defects, correction, mobile checks, publishing, and proposed updates.

Independent verification, editorial review, reader comprehension, accessibility, later refreshes, and civic use.

Source-to-claim trace, editorial gate, delivered desktop/mobile/access states, audience test, correction policy, and update owner.

Governed BI and executive use

Product contracts expose semantic models, queries, permissions, and review surfaces; testimony explains the need for narrow, owned metrics.

Routine task success, review queues, refusal, one-number incidents, trust recovery, and total cost in one independent study.

Approved metric contract, exact query and filters, permissions, owner, refusal behavior, escalation, and decision outcome.

Learning and explorable explanation

One 117-person experiment measures immediate comprehension; one classroom case shows direct instruction rescuing failed AI-assisted construction.

Delayed construction, unfamiliar transfer, ordinary assistant use, and a complete learner-to-reader project.

Immediate and delayed unassisted performance, diagnosis, transfer, confidence calibration, and the eventual audience’s comprehension.

Accessible and small-screen reading

Controlled adjacent studies measure low-vision smartphone and blind nonvisual outcomes; other sources expose generated-description failure and newsroom scale.

AI-assisted authoring joined to disabled readers on their own devices, with verification, recovery, and harm measured.

Co-designed alternatives, real assistive technology, success, error, time, workload, independent verification, feedback, and remediation.

A practical evaluation

Measure the complete creator-to-artifact-to-reader episode.

A prettier first render is not enough evidence for a workflow change.

  1. 01

    Name the environment and audience.

    Exploration, governed BI, operations, public explanation, science, education, or a reusable application carry different obligations.

  2. 02

    Use representative work.

    Include self-owned data, awkward semantics, a correction request, and the actual delivery surface. Retain a simple task as a control.

  3. 03

    Record the current baseline.

    Measure human time, errors, correction path, and delivery effort before claiming a gain.

  4. 04

    Separate first draft from acceptance.

    Count prompting, waiting, cleanup, verification, publishing, and monetary cost.

  5. 05

    Inspect consequential choices.

    Make fields, filters, aggregation, transformations, code or semantic query, and generated interaction state visible.

  6. 06

    Test the delivered artifact.

    Recompute values, exercise controls, check exports, accessibility, responsive behavior, and future maintenance.

  7. 07

    Test readers separately.

    Ask defined readers for the claim, evidence, uncertainty, and next action. Measure correctness, time, confidence, and harmful misreadings.

The updated gap ledger

Correction, reader harm, accessibility, and one handoff now have bounded evidence. Production reliability, maintenance rates, and total cost do not.

“Partly answered” means one bounded study exists—not that the field can generalize.

Partly answered
  • Correction and premature acceptance108 engineered-error episodes; no field distribution
  • Reader comprehension and harm117-person scaffolding study + 48-person misleading-chart study
  • Mobile and nonvisual outcomescontrolled non-AI delivery studies
  • AI-assisted BLV learningpreference and spatial model; no accuracy lift
  • Second maintenanceone direct episode; no rate or comparison
Still materially open
  • Field correction and abandonment distributions
  • Local semantic accuracy
  • Authenticated production across contexts
  • Maintenance and regression rates
  • Total human + model cost
  • Accessible delivery on readers’ own devices
  • Decision quality and trust calibration

How to read the evidence

Papers, product pages, and public testimony do different work.

Measured study

Behavior under defined conditions

Supports claims about its participants, tasks, models, measures, and comparisons. Small or dated studies do not automatically transfer to current production.

Practitioner survey

Reported use in a defined respondent sample

Shows which jobs and tools respondents report. Self-selection, missing items, and sponsor or community channels limit population claims.

Product documentation

Current interaction contract

Shows what the provider says the feature can see and do. It does not independently establish correctness, usability, or adoption.

Public testimony

An experience that occurs

Surfaces jobs and failure costs that experiments may omit. Self-selection means it cannot establish prevalence or comparative performance.

Evidence cut 14 August 2026. Twenty-one human studies and structured workplace evaluations were read in full alongside practitioner surveys, current provider and open-source documentation, public discussions, first-person production cases, and dated essays and interviews from several professional positions. The strongest counter-reading is that better models will erase older failures. Fresh evidence shows real improvement in visual output—and new failure surfaces from richer, slower, more complex artifacts. Capability is moving; the location of human work is moving with it.

Publication history

Update log

  1. Initial public edition mapping practitioner jobs, tool choices, creator experience, delivery failures, reader reactions, and open evidence gaps.

See updates across the research package