The state of AI-assisted data visualization Markdown source

Research snapshot · Evidence reviewed through August 14, 2026

The practitioner and reader experience of AI-assisted data visualization

Status: research snapshot, evidence cut 2026-08-14. Recheck by 2026-11-14, or earlier after a material change in the major general-purpose, visualization, or business-intelligence assistants discussed below.

Companions: The state of AI-assisted data visualization research, Human skills and banked gains in AI-assisted data visualization, and Learning data visualization when AI can make the chart.

This report asks a different question from the research review. Instead of starting with architectures and benchmarks, it asks what it is like to use AI to make or read a data visualization in August 2026. It covers the jobs people are trying to perform, the kinds of tools available, measured studies of human behavior, and public accounts of what feels remarkable or frustrating.

The short answer is that AI is already good at getting someone from a blank page to a plausible first thing. It is much less reliable at carrying the work through local meaning, correction, delivery, and audience understanding. That distinction explains why a demo can be astonishing, a practitioner can save real time, and the finished chart can still be wrong or unhelpful.

Executive summary

AI-assisted visualization has three different success thresholds:

  1. Possible: a model or tool can produce a chart, dashboard, analysis, or interactive artifact under some conditions.
  2. Accomplished: a practitioner can turn that output into correct, inspectable, editable, and deliverable work at an acceptable total cost.
  3. Experienced: creators understand and can control the process, while readers understand the result, notice its limits, and place appropriate trust in it.

Most public demonstrations stop at the first threshold. The strongest research systems reach into the second by exposing transformed data, code, semantic models, edit history, or direct controls. Very little work reaches the third: independent reader comprehension, decision quality, calibrated trust, accessibility, and mobile use remain sparsely measured.

Across the evidence, five conclusions are reasonably stable:

A sixth conclusion is becoming visible but is less settled: adoption is entering through existing work rather than replacing the visualization stack. In the best available practitioner survey, AI use clustered around preparation, analysis, coding, and ideation, while Excel remained the most widely used visualization tool. Current BI testimony similarly describes chat taking ad hoc questions while recurring dashboards retain shared definitions. This is a directional reading of self-selected evidence, not a market-share estimate.

The practical implication is not “use AI” or “avoid AI.” Use it where the work is reversible, the output remains inspectable, and the human can recognize a bad result. Demand stronger grounding, deterministic checks, and explicit reader testing as the semantic stakes, delivery costs, or audience consequences rise.

The gap that organizes the evidence

The same output looks different at each threshold. A generated dashboard can render and still fail as work; a correct chart can still fail its reader.

Threshold The question being answered Evidence that counts Common false substitute
Possible Can the system produce the requested class of output at least sometimes? executable code, a rendered chart, task completion, supported feature contract a polished screenshot or vendor description
Accomplished Can someone finish the actual job correctly and maintain the result? correct data and transformations, inspectable decisions, correction cost, delivery, reuse, total time and cost first-pass speed, nonblank rendering, or a benchmark score alone
Experienced — creator Can the person understand, steer, correct, and appropriately trust the process? observed behavior, successful repairs, calibrated confidence, workload, abandonment, learning satisfaction or stated preference alone
Experienced — reader Does the final audience understand the claim and its limits and make an appropriate judgment? comprehension, recall, decision quality, trust calibration, accessibility, behavior on the delivered device model chart-reading, aesthetic ratings, or author confidence

This is not a maturity ladder on which every visualization must climb. A throwaway sketch may only need to be possible. A board metric, public-health chart, or published explanatory graphic needs a defensible path through all four rows.

Who is actually reaching for what

The evidence is better at showing where AI has entered the workflow than at naming a winning product. The 2024 Data Visualization Society survey was self-selected but unusually useful: 980 practitioners started it, 763 completed it, and 825 answered the AI-use question. Of those 825, 305 (37.0%) said they had used AI in visualization work during the prior year, up 13 percentage points from 2023; 492 (59.6%) said no and 28 (3.4%) were unsure.

Among the 305 AI users, multiple selections were allowed. The work was weighted toward the beginning of the lifecycle:

Job reported by AI users Respondents Share of 305 AI users
Prepare or clean data 184 60.3%
Analyze data 134 43.9%
Ideate or storyboard 119 39.0%
Produce visualizations 92 30.2%
Support a visualization team 27 8.9%

This is not a story of visualization specialists abandoning incumbent tools. Excel was used often or sometimes by almost identical shares of AI users and nonusers in the public microdata (66.9% versus 66.7%), and Tableau use was also similar (43.9% versus 47.4%). AI users were more likely to report Python (42.6% versus 27.4%), D3 (30.5% versus 13.4%), Observable (16.1% versus 5.7%), Figma (40.0% versus 28.5%), and Datawrapper (19.7% versus 10.2%). The defensible interpretation is that early adopters in this sample often already crossed code, design, and publishing environments. Co-use does not show which tool caused adoption.

A second, non-visualization-specific survey sharpens the split. In the 2025 State of Analytics Engineering, 80% of 459 respondents reported day-to-day AI use, up from 30% one year earlier. Seventy percent used AI for analytics development, mainly through general assistants such as ChatGPT, Claude, and Gemini; roughly 25% used specialized AI inside development tooling. Only 30% were using natural language to consume data, while another 29% wanted to and 23% had experimented. The sample came from a vendor community and is not visualization-specific, but it reinforces a plausible sequence: code and documentation adoption precede trusted conversational consumption.

The audience-to-environment map

The table separates four kinds of evidence that are often blurred together: surveyed behavior, public cases, provider-defined target users, and our inferred fit from the work surface.

Who What they appear to reach for first Job they are trying to do Status of the evidence
Visualization practitioners with an existing stack a general assistant beside Excel, Tableau, Python, R, D3, Figma, or a publishing tool clean data, debug code, learn a technique, storyboard, draft labels, make a first view Observed survey pattern. The DVS sample measures AI jobs and incumbent-tool co-use, not specific AI-product share.
Analytics engineers and code-capable analysts ChatGPT, Claude, Gemini, Copilot-style code help, then project-aware notebook or development agents write and explain SQL or Python, document models, debug pipelines, generate a chart inside an analysis Observed survey pattern. The dbt sample is vendor-adjacent and broader than visualization.
Spreadsheet-native analysts, operations teams, and finance users AI inside Excel or Sheets; newer AI spreadsheets when code or live connections outgrow a grid ask about a range, create formulas and charts, preserve familiar cells, hand work to colleagues Installed-base inference plus provider target. Excel is widely observed; adoption of its AI features was not measured here.
BI authors and data teams AI inside Power BI, Tableau, Looker, Hex, Sigma, or a governed data platform author reports, define or reuse measures, inspect queries, answer stakeholder follow-ups, govern access Provider target plus public cases. Independent head-to-head field evidence remains absent.
Business consumers and executives conversational layers such as Genie, Spotter, Ask Sigma, Qlik Answers, Amazon Q, or Oracle’s assistant retrieve a metric, ask why it changed, get an ad hoc cut, avoid navigating a complex report Provider target with self-selected testimony. Current cases say chat complements recurring dashboards and loses trust quickly without curated definitions.
Designers, journalists, and explanatory communicators general assistants for ideation or code, then Figma, Datawrapper, Flourish, or a newsroom-owned production stack explore, create variants, annotate, improve accessibility, implement a story without surrendering editorial judgment Survey subgroup and workflow inference. Journalism had 24 AI users among 47 valid DVS responses, too small and self-selected for a population rate.
Developers and data-app builders coding agents, generated applications, declarative chart specifications, open-source agent frameworks, and reusable skills make bespoke interaction, integrate data and UI, automate repeatable production, deploy an application Project positioning and inferred fit. Repository activity shows availability, not routine user success.

The July 2026 public BI discussion helps explain the business-consumer row without estimating its size. Reported deployments use chat for exploration, follow-up questions, and requests that previously became analyst emails; recurring dashboards still provide shared monitoring and durable definitions. Participants repeatedly tied trust to a narrow mart, approved metrics, visible refusals, and an owner for each measure. Several described slow or failed pilots, and one reported that a single wrong executive number ended trust. These are contemporary cases, not prevalence evidence.

The jobs: visualization is a workflow, not a prompt

“Make a chart” hides work before, during, and after visual encoding. Separating the jobs matters because the same tool can be strong at one and harmful at the next.

Phase Job Useful assistance Human responsibility that does not disappear
Define Frame the question restate a request, propose hypotheses, identify missing information decide what matters, tie the work to a decision, reject an incoherent request
Define Acquire and govern find likely tables, draft access requests, describe schemas choose authoritative sources, respect permissions, establish ownership and allowed use
Work the data Prepare and transform clean types, join, reshape, write SQL or Python, create routine calculations define measures, resolve ambiguous fields, inspect exclusions and missingness
Work the data Analyze calculate summaries, fit standard models, compare groups, expose anomalies choose a valid method, interpret uncertainty, know which causal claim is unavailable
Work the data Explore generate many cheap views, suggest cuts, translate follow-ups into operations recognize spurious patterns, investigate causes, know when to stop branching
Make the artifact Sketch or reproduce remove blank-page work, copy a reference, create dashboard mockups and variants decide whether the reference belongs in this context and what must change
Make the artifact Choose and encode suggest chart forms, mappings, aggregations, annotations, and emphasis make the analytical and rhetorical choice for this data, claim, and audience
Make the artifact Critique and refine flag common perceptual problems, apply repetitive settings, draft labels and accessibility text resolve composition, tone, brand, direct-manipulation details, and conflicting advice
Make the artifact Verify and debug expose code and intermediate tables, generate tests, replay interactions, compare results state expected behavior, choose decisive checks, judge local semantic correctness
Deliver Package and publish assemble a report or app, write documentation, export standard formats accept the delivered artifact, satisfy privacy and review gates, test the actual surface
Deliver Explain and interrogate provide tooltips, summaries, guided questions, and grounded conversational follow-ups preserve the source claim, communicate uncertainty, evaluate reader understanding
Sustain Maintain and update regenerate routine artifacts, flag drift, propose regression checks own dependency changes, semantic revisions, handoff, correction, and retirement

The public cases are most consistent about work adjacent to the final visual: SQL help, calculations, documentation, mockups, tooltips, bulk changes, and alternative ideas. Claims that a system “built the dashboard” often compress substantial context: an existing screenshot, prepared source tables, a semantic model, a known template, or extensive cleanup.

Listening to the field: what this work actually feels like

The accounts below are selected, not sampled. They come from first-person project logs, interviews, public workplace discussions, and one qualitative study of biomedical-visualization practitioners. Public handles, roles, and outcomes are self-described unless a named publication or study says otherwise. The accounts demonstrate that an experience was reported; they do not estimate how common it is or prove that the artifact worked as described.

One distinction makes the listening more useful: attitude and behavior are not the same dimension. A person can be delighted and unwilling to ship, anxious and still using the tool, skeptical but grateful for boilerplate help, or enthusiastic while retaining direct inspection and repair. A 17-person biomedical-visualization interview study observed five such modes, from enthusiastic adoption through skeptical avoidance. That specialized typology is not imposed on every account here. It is a reminder not to reduce the field to boosters and opponents.

Before the work: invitation, pressure, and hesitation

The first candidate: access, velocity, and the rush of possibility

The working loop: babysitting, interruption, and selective handback

Acceptance: reputation, accountability, access, and the right to refuse

After delivery: use, learning, identity, and whether the saved effort mattered

Listening past the original creator

A second listening pass pursued people who were mostly absent above: routine spreadsheet workers, quiet non-users, recipients of generated analysis, and the people expected to live with the result. It changes the picture in ways a collection of creator success stories cannot.

Setting and voice Reported episode What it makes visible Limit
Public-sector staff using Microsoft 365 in ordinary work A DWP mixed-methods evaluation found that people routed work according to expertise, time, trust, data sensitivity, and habit. Experienced or data-heavy work often stayed in established methods; Excel users specifically reported trouble with complex data, charts, and formatting. One person valued the assistant yet was too busy to remember to use it consistently. Non-use can be local and quiet. A person may value the overall assistant while declining it for the spreadsheet, chart, or stakeholder-owned task in front of them. 1,716 user survey responses, 2,535 comparison responses, and 19 interviews provide substantial workplace evidence, but licences were not randomized, no pre-trial baseline existed, outcomes were self-reported, and visualization is a subset.
A supply-chain analyst and a spreadsheet-help community A self-described intermediate Excel user said company encouragement made AI the first stop for reports and reduced visits to a peer forum. The 190-entry discussion ranged from useful Power Query help to invented functions, wrong references, damaged formulas, and decisions never to ask again. Several people worried about losing the learning that comes from solving another person’s problem in public. Assistance changes the support ecology, not only task time. Instant private answers can lower access cost while thinning public explanation, correction, and incidental learning. One self-selected thread establishes that these experiences and concerns were reported. It cannot establish prevalence or a forum decline caused by AI.
People making simulated industrial decisions Twenty participants used both a dashboard and an LLM conversational interface. They experienced conversation as compressed retrieval and mental offloading, but often regarded the result as provisionally useful rather than defensible. When speed was the criterion, 15 chose the chatbot and 3 the dashboard; when confidence was the criterion, all 18 decisive responses chose the dashboard. Several wanted both: overview to discover the question, conversation to retrieve or assemble the answer. Usefulness and reliance separate for recipients just as delight and acceptance separate for creators. Extra visual interaction can be valuable when it exposes the values and relationships needed to defend a decision. The exploratory study used four short simulated tasks; 18 of 20 participants had computer-science backgrounds and were proxies rather than experienced industrial decision makers.
Program recipients and an observer of an executive sales dashboard A social-impact dashboard case reports that program managers, evaluators, and funders wanted overview and drill-down, definitions, neutral language, comparative context, and review support. In a separate named public critique, telling an executive to verify regenerated charts by checking raw data felt like transferring the product’s verification burden to the person seeking a quick decision. A recipient does not simply need a shorter summary. The path from alert to explanation and from claim to evidence must be proportionate to the recipient’s actual job. The design case omits its interview count and instruments; the critique is one observer’s account and does not establish that the output was wrong or how the intended executive reacted.
Blind and low-vision people learning unfamiliar chart forms In a 12-participant study, people learned violin plots and clustered heatmaps with either tactile charts plus text and GPT-5.2 or text and GPT-5.2. Eleven preferred the tactile combination; one said the better mode depended on chart complexity; none preferred text and chat alone. Participants described touch as supplying the spatial model and chat as supplying flexible clarification. Accessibility is not one generated description. Complementary modalities can answer different needs, and an open-ended assistant cannot always detect misunderstanding or supply questions a novice does not know to ask. Preference and reported spatial understanding improved, but measured chart-understanding accuracy did not. This was a short study with a small English-speaking US sample, not a long-term learning or delivered-work evaluation.
A second maintainer inheriting an AI-built dashboard A public discussion begins with a dashboard failing while its creator was away. The inheriting maintainer found that refresh and scheduling depended on the creator’s laptop and replaced the setup with a proper pipeline. Replies described missing source and metric semantics, hidden dependencies, absent comments and documentation, and an implicit service expectation nobody had owned. The maintenance boundary begins before handoff. Execution host, refresh lineage, definitions, dependencies, documentation, ownership, and rollback must be inspectable while the creator is still present. This is one self-selected, unverified public episode plus discussion. It establishes that the experience was reported, not how common it is or whether AI rather than ordinary software practice caused every defect.

Quiet withdrawal is therefore a gradient: task-level refusal, habitual fallback, feature abandonment, substitution of one model for another, and provisional use without confident reliance are different behaviors. A retained licence or active session can hide all five.

The direct second-maintainer gap is no longer empty, but it remains extremely thin. One episode is enough to make the hidden execution host and ownership boundary concrete; it is not enough to estimate frequency, compare AI-assisted and direct work, or describe what happens across organizations and tool stacks.

What these voices reveal that aggregate capability evidence does not

What people in the field say has changed

Benchmarks tell us whether a system can perform a defined task. The perspectives below explain why that task matters, what surrounds it in practice, and where people disagree about the purpose of assistance. They include first-person build diaries, practitioner essays, interviews, design-studio reflections, a classroom exercise, and two research papers. They establish that an experience or argument exists. Except where a measured study is named, they do not establish frequency, comparative performance, or a general causal effect.

Across these selected perspectives, implementation is getting cheaper. The substantive disagreement is what the saved effort should buy.

Tension What one perspective contributes What the counterweight contributes What context should decide
Construction speed versus verification debt Two 2026 newsroom diaries describe work that once took weeks appearing in days or hours. They also report incorrect totals, misplaced geographies, missed notes, repeated correction, usage limits, mobile checks, and new validation machinery. Analytics practitioner Benn Stancil argues that a plausible chart cannot validate the calculation behind it; local definitions, executable lineage, and review remain necessary. Semantic stakes, independence of the check, cost of a wrong result, and whether the complete pipeline—not only the first render—is owned.
Lower implementation barriers versus more valuable judgment Enrico Bertini maps opportunities across acquisition, wrangling, analysis, data tours, and creative exploration. Alberto Cairo treats easier code as capacity extension. Both make question quality, synthesis, interpretation, and final analytical choice more important rather than less. Cairo also distinguishes exploratory, explanatory, essayistic, artistic, and poetic visualization purposes. Purpose, audience, consequence, the person’s capability profile, and whether the resulting claim can be inspected and challenged.
Frictionless production versus productive friction Removing syntax and formatting work can let a person reach a candidate before the question goes cold. In one classroom exercise, students spent 25–30 minutes failing with AI before direct instruction supplied the missing chart structure. An analog data-journalism reflection describes slow, manual work as a source of attention, experimentation, and care. Whether the effort is obsolete mechanics or practice in structure, skepticism, memory, and meaning; whether assistance will later be withdrawn.
Automated access versus independent verification Generative systems can draft descriptions and alternative representations at a scale that newsrooms have struggled to support. Elavsky and Xiong Bearfield show how model-to-model chart descriptions can lose data, source, uncertainty, design choices, and authorial purpose. Blind journalist Johny Cassidy describes access as an organizational, multimodal, and reader-feedback problem—not an alt-text checkbox. Whether the reader can test the account, reach the underlying data or authorial description, choose another modality, report a failure, and obtain remediation.
Static charts versus conversational or divergent representations Richard Brath asks which visualizations remain useful when a model can extract and organize insights. Elijah Meeks proposes making audience, intent, suggestion, interpretation, and conversation first-class framework objects. Domestic Data Streamers argues that visualization remains a distinct language for comparison, human experience, and creative divergence that a textual answer does not replace. Reader task, need for overview or a shared evidence surface, qualitative meaning, desired novelty, and whether interaction leaves an auditable trace.

An early formal baseline helps distinguish current observation from hindsight. In Doom or Deliciousness, 21 visualization, HCI, art, and machine-learning experts interviewed in 2023 anticipated rapid prototyping, personalization, and creative assistance alongside bias, unreliable output, inattention, lost agency, and unearned trust. The study used a convenience sample, foregrounded image-generation systems, and asked participants to think beyond the technology they had used. It is useful as a “what aged well?” comparison, not a verdict on current systems.

Two current build diaries show where the time moved

The two journalism cases are unusually detailed accounts of current production. They are still self-reports by the people who built and checked their own work. Keeping the gain beside the incurred work prevents “built in a week” from becoming a quality claim.

Lifecycle stage Ottaviani: two reconstructed projects Tordecilla: national health dashboard What neither case establishes
Baseline and first candidate One map was rebuilt in two days after taking roughly three weeks in 2012; a second portal was attempted in a much more autonomous pass. A repository became a recognizable site in about an hour; the full dashboard, maps, charts, analysis, and checking workflow took one week. An equal-budget comparison with the best current direct workflow or another practitioner.
Defects and correction Incorrect cluster and chart totals and incomplete bilingual text required source comparison and several correction rounds. Highly urbanized cities rendered in the sea or in broken shapes; source notes and exact evidence were sometimes missed; long runs took shortcuts. A field distribution of error, correction time, regression, premature acceptance, or abandonment.
Architecture and control The author decomposed the work into about 20 tasks, ordered foundational work first, and kept code in an open repository. The author separated extraction, data, presentation, insights, and fact-checking so a source error could be repaired and replayed. That decomposition or a model-built validator independently improves correctness across projects.
Delivery and readers The projects were published as working prototypes; the author identifies attention, understanding, and civic use as the remaining bottleneck. Mobile rendering was checked and stakeholders reacted to the tool, but no reader-comprehension or decision study was conducted. Audience understanding, accessibility, consequential use, or sustained adoption.
Sustainment Proposed next steps include record-by-record reconciliation and scheduled updates. Refactoring and longer-term maintenance remained after the one-week build. Later refreshes, dependency changes, second-person handoff, incident response, or retirement.

The tools: a showcase of the current approaches

There is no single AI-visualization product category. A useful comparison asks five literal questions: What can the assistant see? What can it change? Which intermediate state can the user inspect? How local is a correction? What form is actually delivered? The examples below are a current landscape, not a ranking. Product pages establish feature contracts and intended users, not independent quality or adoption.

General assistants and specialist file chat

Tool Who it is positioned for What it makes or changes Where control returns to the person
ChatGPT data analysis anyone bringing a bounded file or table into a conversation cleaning, executed analysis, common static or interactive charts, explanations, downloadable results inspect generated code and tables; supply local definitions and move production work elsewhere
Claude Artifacts people who want a shareable custom visual or small application generated HTML or application code, interaction, explanatory interfaces inspect and edit the artifact; assume responsibility for state, accessibility, hosting, and maintenance
Julius file- or connection-oriented analysts seeking a dedicated data chat analysis, charts, statistical models, reports, and presentations continue through prompts or exported artifacts; independent repair and correctness evidence is sparse

Spreadsheet-native assistance

Tool Who it is positioned for What it makes or changes Important contract boundary
Copilot in Excel Excel beginners through advanced analysts Python-backed answers and optional static chart or table insertion; an advanced mode can create refreshable Python cells the direct-answer mode does not modify the workbook; inserted charts are static and non-refreshable
Gemini in Google Sheets people already working in a shared spreadsheet summaries, formulas, prompted edits, and editable charts inserted with supporting data on a new tab the generated chart links to the new tab, not to subsequent changes in the original dataset
Bricks teams wanting spreadsheets, dashboards, slides, and collaboration in one workspace grid-based analysis, live dashboard boards, presentation material combines surfaces to reduce handoff, but current evidence is provider documentation rather than comparative use
Quadratic spreadsheet users who also need Python, SQL, JavaScript, or live database connections AI-generated code and queries, direct grid edits, code cells returning results to the sheet exposes code and schema inside a familiar grid; asks users to manage a more technical workbook

Notebooks and visible canvases

Tool Who it is positioned for What it makes or changes Important contract boundary
Hex AI technical notebook authors, data teams, and consumers of published data apps SQL, Python, chart, pivot, and Markdown cells; generated apps; semantic models; consumer chat roles are explicit: technical users audit code, consumers query curated apps, managers govern semantic context
Observable Canvases AI collaborative analysts who want AI work beside their own new tables, transformations, SQL, and charts placed on the current canvas the assistant adds new versions rather than silently editing or deleting; its view is bounded by viewport and schema

Guided storytelling editor

Tool Who it is positioned for What it makes or changes Important contract boundary
Flourish AI beginners, visualization practitioners, and teams refining a chart in a known editor native editor settings for styling, labels, sources, annotation, and accessibility proposed changes remain reversible in the editor; the assistant does not edit source data or invent missing intent

Governed BI and data-cloud assistants

Tool Primary surface Intended reach Distinguishing control or dependency
Power BI Copilot report authoring inside Power BI BI authors creating or modifying report pages grounded in the model and report state; quality depends on maintained measures, descriptions, permissions, and review
Tableau Agent Tableau web authoring analysts building views and calculations in an existing workbook works through the worksheet and data source, retaining a conventional authoring surface for correction
Looker Conversational Analytics natural-language questions over LookML business users querying governed fields and measures administrators own glossaries, defaults, verified queries, and semantic-model quality
ThoughtSpot Spotter search and conversational analytics business teams asking everyday or strategic questions provider positions data teams as maintainers of governed models rather than answer queues
Ask Sigma conversational analysis over a cloud data workspace non-specialists and analysts exploring organization data shows data sources, formulas, filters, and analysis steps and allows step-level edits instead of full regeneration
Qlik Answers structured and unstructured enterprise knowledge business users asking questions and authors assembling analytic output combines RAG, chart and dashboard agents, permissions, and audit logs; evidence remains provider-authored
Databricks Genie a curated, no-code data chat business users asking questions beyond existing dashboards returns SQL-backed answers and visualizations; authors monitor, review, and refine instructions and definitions
Amazon Q in QuickSight BI authoring, Q&A, executive summaries, and data stories authors and consumers already in the AWS BI surface integration reduces handoff; correctness still inherits datasets, topics, permissions, and author review
Oracle Analytics AI Assistant natural-language analytics in Oracle authors who prepare metadata and consumers who request charts or narratives explicitly separates the author who prepares data and metadata from the consumer who asks questions

Mixed-initiative and open-source systems

Project Approach Why it matters Evidence boundary
Data Formulator direct encoding controls plus natural-language transformation, branching, visible derived data, and report sharing treats AI as one control surface in an inspectable analysis environment research evaluations are small; the current repository is more capable than the studied GPT-3.5 system
Lumen declarative data pipelines, visualizations, dashboards, and specialist agents that can be serialized and edited makes generated transformations and output inspectable and reusable across chat, notebooks, and dashboards open-source feature surface, not independent evidence of routine field success
Vizro low-code Python dashboard specification with code escape hatches and agent-facing tools constrains generation to a production-oriented component system while preserving extension and deployment paths a reusable framework reduces some errors; acceptance still requires tests and design judgment
AntV chart visualization skills retrievable, version-specific instructions and declarative chart or code generators for several visualization libraries shows how a harness can supply library syntax, chart vocabulary, and guardrails at generation time the project reports its own 174-case testing; independent and current-model ablations are still needed

Six techniques now cut across product classes: unrestricted code generation; semantic grounding in governed fields and measures; structured or declarative intermediate representations; direct manipulation and local undo; deterministic render or data checks; and proactive reader scaffolding. These techniques can coexist. The meaningful choice is which one contains the most expensive failure for the actual job.

A practitioner may therefore ask a general model for dashboard concepts, use a coding agent to generate measures, edit the layout manually in BI software, and let a reader-facing assistant explain the final report. That is one workflow with four evidence and control boundaries—not one assistant experience.

What feels amazing to creators

A first draft appears before the idea has cooled

The most consistent delight is immediacy. A person can move from a question or rough picture to something concrete enough to react to. In the DashChat study, participants valued rapid multi-chart mockups and simulated data as a way to negotiate dashboard requirements before production data existed. Data Formulator participants could combine direct encoding choices with short transformation prompts and complete all 16 reproduction tasks, although the eight-person study did not include a conventional-tool baseline.

This is valuable because a draft changes the conversation. “Show me three ways to compare this” is easier to critique than an abstract discussion of a future dashboard. The benefit is not that the first output is right; it is that the cost of making a candidate has fallen.

Tedious, legible work compresses well

Practitioners repeatedly describe value in SQL, calculations, documentation, bulk renaming, standard chart setup, tooltips, formatting, and alternative approaches. These tasks have inspectable outputs and a relatively short path to correction. In one public BI discussion, contributors described using AI to explain dashboard queries in documentation, prepare stakeholder questions, review junior work, critique mockups, and accelerate unfamiliar API or Python tasks. This is public testimony, not a survey, but it gives the broad phrase “AI in BI” concrete content.

AI can bridge representations a person does not fluently command

The 22-analyst verification study found that people moved among natural-language explanations, code, original data, intermediate data, results, and visual summaries. When Python was unfamiliar, some participants translated the logic into SQL or operations they knew. AI is useful here as a representation bridge: it can express the same intent as prose, code, a table, or a visual candidate. The person still needs at least one representation in which they can evaluate the result.

A bounded integrated tool can be genuinely faster

In a randomized public-health analysis exercise with 30 analyzed participants, the integrated ChatGPT environment finished faster than the distributed R/Stata-plus-ChatGPT workflow: a median 38 versus 45 minutes. The integrated group’s visualizations were also judged more meaningful and correctly labeled. That is real evidence for reducing tool-switching and syntax work in this specific 45-minute exercise. It is not evidence that the integrated group’s overall work was safer or better; that result points in the other direction.

A critic can improve work without becoming the designer

Visualizationary followed 13 designers—six novice, four intermediate, and three expert—across a three-to-five-day window as they iterated at least five versions of a visualization made with their own data and preferred tool. Its GPT-3.5 critique was combined with deterministic perceptual filters, a hierarchical report, and version tracking. Three external experts later rated the final work as improved on average (3.69 on a five-point scale where three meant no change).

The system did not supply a universally correct edit list. Participants used 68 pieces of feedback, ignored advice that conflicted with intent or seemed wrong, and did not always end with their best-rated version. Intermediate and expert designers generally converted critique into better revisions more effectively than novices. The small, short study has no conventional-critique baseline, but it demonstrates a credible role for AI as inspectable counsel inside an existing practice rather than as an autonomous author.

What makes the work hard

The model can make an invisible analytical decision

Aggregation, filtering, grouping, baseline, and scale choices can be buried in generated code or made implicitly from an ambiguous prompt. In the 2026 novice study, 13 of 20 participants issued ambiguous or problematic prompts, and the system sometimes chose inconsistent aggregations or supplied interpretations the user could not verify. A chart can look conventional while encoding a different question from the one its creator thinks it answers.

Governed BI environments try to narrow this gap through semantic models, descriptions, verified queries, permissions, and authored instructions. The current Looker documentation is unusually explicit: the model maps natural language onto maintained LookML fields and lets administrators supply business glossaries, defaults, and verified queries. That is an architecture contract, not proof of perfect answers. It also makes the organizational dependency visible: conversational analytics inherits the quality of the semantic model someone maintains.

Correction can be harder than creation

Natural language is a powerful coarse control and often a poor fine control. In the novice study, 11 of 16 attempts to fix clutter and eight of nine attempts to repair unusable charts failed. More capable systems generated richer interfaces, but also introduced longer waits, unclear selected states, broken controls, and cases where a model interpreted a blank rendering as if it contained data.

Public Power BI accounts add a practical version of the same problem. In a 35-entry thread, contributors described useful greenfield structure alongside stale filters, bookmark IDs, whitespace cleanup, slow preview cycles, and minor changes that were faster to make manually. One contributor summarized the division as trusting and verifying generated code while retaining human judgment over whether business users would accept the report. These are cases, not frequencies.

Another practitioner reported a generated inline calculation that produced a desired bar chart but could not be located or edited through familiar Power BI mechanisms. The issue is not merely whether the number happened to be right. The generated implementation broke the user’s ability to inspect, maintain, and explain it.

The longitudinal Visualizationary study found a quieter form of the same cost: feedback could be verbose, vague, or based on a mistaken reading of intent. Novices in particular struggled to translate broad advice into a precise edit, and the final version was not always the highest-rated one. A critique system therefore needs triage, rationale, and edit locality; more feedback is not the same as a better correction path.

A 2024 controlled study of AI-assisted data analysis gives this problem an actual, if deliberately artificial, denominator. Eighteen experienced analysts completed 108 task episodes designed so that the model would make two errors unless the person found and repaired them. Seven episodes were not completed within 15 minutes. In 31 episodes—28.7% of the study tasks—the participant said the work was complete while an issue remained. Interfaces that exposed editable assumptions and plans improved participants’ sense of control, but the study did not detect a difference among conditions in task success, completion time, or verification hints.

This is evidence that premature acceptance and non-completion occur even among experienced analysts who have been warned to look for errors. It is not a field failure rate: every task was engineered to be troublesome, the researcher supplied correctness checks and hints, and no episode continued through publication or abandonment in ordinary work.

Total cost includes waiting, prompting, cleanup, and model usage

First-pass generation time omits the rest of the loop. Practitioners mention prompt construction, regeneration latency, application refreshes, token or credit consumption, deployment cycles, and the cognitive cost of deciding whether a strange result is a model error or their own misunderstanding. A one-minute generation can be slower than a ten-second direct edit; a 25-minute dashboard rebuild can be impressive and still depend on an existing screenshot, prepared tables, and substantial layout review.

The useful unit is therefore total human-plus-system time to an accepted artifact, including correction and verification—not seconds to first render.

The same 108-episode study illustrates why interface improvements should not be translated automatically into time savings. Mean completion time was 543 seconds for the conversational baseline, 588 seconds for stepwise decomposition, and 658 seconds for phasewise decomposition; the differences were not statistically significant. More visible structure made correction feel more controllable while also creating intervention and information costs. The study measured human task time, not model inference, subscription cost, later cleanup, or maintenance, so it still does not provide a total-cost ledger.

Production adds failure surfaces that studies often remove

Controlled studies usually provide clean data, bounded tasks, and a working prototype. Real work adds permissions, confidential data, incomplete semantics, unreliable connectors, browser behavior, responsive layout, organizational review, publishing systems, and future maintenance. The 2025 randomized exercise recorded upload and browser problems even in a simulated setting. Practitioners in public threads describe network policy, application refresh, deployment, and cost constraints that are normally absent from benchmark scores.

Interviews with 17 biomedical-visualization practitioners show where a high-stakes production boundary is currently being drawn. Thirteen already used generative AI somewhere in their workflow, primarily for auxiliary work such as research, moodboards, boilerplate code, background assets, captions, or translation. All opposed incorporating substantial AI-generated visual content into final scientific work, citing validation, accuracy, copyright, provenance, and accountability. That is evidence about real production and dissemination judgment, not delivery efficacy: the study did not observe projects through handoff, publication, reader use, refresh, or maintenance.

Speed and quality separate in the measured studies

The most cautionary evidence is not that systems always fail. It is that their success signals can disagree.

Study What improved or looked good What did not follow
Randomized public-health analysis exercise, 2025 — 30 analyzed participants Integrated AI users finished seven minutes faster at the median; their charts were judged more meaningful and correctly labeled. Overall scores were not significantly different. Only 6.7% of integrated submissions versus 26% of distributed-tool submissions were free of serious errors. Small, underpowered study; no non-AI control.
Vibe Visualizing, 2026 — 20 novices, 60 sessions, 175 charts Participants could repeatedly produce charts and reported mean confidence of 3.73/5 and satisfaction of 3.93/5. Replayed prompts showed better visual output from newer models on some measures. Every chart had at least one design flaw; 52 of 60 sessions had fatal task noncompliance; 22 charts were unusable; incorrect insights appeared in 12 sessions. Verification was rare and many repair attempts failed.
Data Formulator 2, CHI 2025 — eight corporate participants All participants completed 16 reproduction charts; direct controls, short prompts, visible transformed data, history, and multiple inspection artifacts supported work. Six of eight still needed hints. The tasks were reproduction, not open analysis; no conventional-tool baseline or self-owned data; trust in simple steps could carry too easily into complex descendants.
DashChat, 2025 preprint — formative interviews plus 28-person evaluation Rapid pre-data dashboard prototypes helped make requirements and layouts discussable; designers rated it easier and more effective than a lightly taught Tableau condition. The comparison favored the familiar purpose-built workflow and measured mockups, not real-data correctness, deployment, or ongoing dashboard use. Precise edits and domain conventions remained difficult.

These results should not be pooled into a universal score. They agree on a more useful point: assistance can improve access, speed, or perceived ease without proportionately improving semantic correctness, error freedom, or delivery.

Expertise is both leverage and burden

The simple story that novices benefit and experts do not is wrong. So is the story that experts automatically get the best results.

Visualization expertise is not one rank. The strongest current literacy framework separates the ability to consume, construct, critique, and connect a visualization to its context. Real work also draws separately on data and statistical knowledge, domain semantics, visual design, implementation, situated judgment, and delivery experience. A person can be a strong business reader and a weak programmer, a domain expert and a novice chart author, or an expert implementer who does not know the local measure definitions. AI removes different friction for each profile.

The human-skills companion therefore uses a stricter definition of a gain. Access, productivity, artifact quality, independent learning, verification, reader outcome, and sustained work are separate. Confidence or an AI-assisted final project does not establish learning; time to a first render does not establish productivity to an accepted artifact.

The learning companion follows that distinction across time. It finds that implementation, translation, and on-demand explanation are becoming easier, while delayed independent construction remains mostly unmeasured. It recommends removing inspectable production friction while preserving prediction, data-to-encoding reasoning, comparison, diagnosis, local repair, retrieval, and unassisted transfer.

Experienced practitioners can filter and redirect

In the comparison between ChatGPT advice and human visualization experts, practitioners valued AI for rapid brainstorming, broad option lists, and a neutral first pass. One experienced participant’s description was roughly: ask for 20 ideas, discard 16, use four. That workflow is valuable because the person can perform the filtering. In the analyst verification study, people with data experience inspected distributions and values, while coders inspected code and others translated logic into familiar operations.

The longitudinal designer study adds direct evidence. Intermediate and expert participants were better able to select among critiques and enact useful changes, while novices needed more help turning high-level feedback into design operations. A 2025 single-company case study reached a compatible but much weaker result: eight interviews produced basic, intermediate, and advanced profiles, then one person from each profile used a prototype for about 30 minutes. The basic user needed initial prompting support and missed some interaction affordances; the advanced user demanded source transparency, encountered a wrong scatterplot, and formulated more intricate corrections. Three probe sessions cannot establish efficacy, but they show why one interface cannot equate “easier” with less visible state for every user.

The same expertise exposes prompt overhead

Human experts were still preferred to ChatGPT for accuracy, helpfulness, reliability, adaptability, context, and actionable advice in the 12-practitioner study. The model produced broad, agreeable recommendations; experts narrowed the scope, built common ground, and raised issues the practitioner had not known to ask about. The model sessions were conducted in 2023, so their raw quality comparison is dated. The contextual lesson remains visible in newer public accounts: experienced people often reserve AI for repetitive or unfamiliar work because describing a familiar small edit can cost more than doing it.

Novices gain access without gaining an error detector

The 2026 novice study is the clearest warning. Participants could produce plausible charts, but verification intent was low, only three verification attempts were observed, and satisfaction remained positive. Some users followed inappropriate model suggestions. Poor visualization and data literacy affected both prompting and the ability to diagnose the result.

This creates an expertise paradox: the people for whom the interface removes the largest access barrier may be the least able to tell when it has quietly answered the wrong question. Better defaults help, but they do not eliminate the need to expose the decision the default made.

The reader receives a claim, not a generation process

Creators experience prompts, waiting, code, and corrections. Readers usually see only the final artifact in a report, article, dashboard, presentation, or social post. They infer care and authority from what is visible.

Trust is assembled from cues

In a 2025 preprint, 37 US participants ranked static charts from news, science, government, and infographic contexts and explained their choices. Clarity was mentioned by 31 of 37 participants; source citation, familiar chart forms, integrity, and visual polish also mattered. Priorities were consistent within many individuals and different across individuals. Infographics polarized viewers: some saw accessibility, others saw promotion or clutter.

The study measured deliberative, self-reported trust—not truth detection, comprehension, or behavior. That distinction is decisive for AI-generated charts. A clean chart can signal professional care even when no one inspected its transformations. A cited source can increase confidence without proving that the chart represented it faithfully.

An AI label does not have one predictable effect

The preregistered 2021 Vis Ex Machina experiment showed identical chart panels with human or algorithm labels to 114 participants. Before the task, 60% preferred human recommendations and 15% preferred algorithmic recommendations; actual choices were approximately even. Data relevance dominated, but some participants followed their prior source preference. An algorithmic label suggested precision or error-freedom to some people and lack of human judgment to others.

That study predates generative visualization and does not tell us whether to label current charts “AI-generated.” It establishes that provenance cues interact with prior beliefs and task context; disclosure is not a universal trust switch.

A model reader is not a human-reader test

In a 2026 study of 60 synthetic charts, three multimodal models matched a narrow designer-intent label more often than 24 human participants. The models tended to enumerate structure and values; people formed trend narratives and were more affected by layout and overlap. The finding is not that models are better chart readers. It is that the metric rewarded a task aligned with machine decoding while human reading pursued a different form of meaning.

Using a model to inspect a chart can catch useful defects. It cannot establish that a journalist’s audience understood the story, a dashboard user noticed an alert, or a student learned the concept.

Reader assistance can guide comprehension, not only answer questions

A randomized experiment with 117 higher-education participants compared three ways to support interpretation of a bar chart, communication network, and ward map: a conventional data story, a passive GenAI agent that answered questions, and a proactive agent that asked educator-authored scaffolding questions and gave feedback. All three groups improved from their unsupported baseline. After the assistance was removed, the proactive group’s median score was 6 of 6, compared with 5 for both the data-story and passive groups; the between-group effect was statistically significant and medium in the authors’ analysis. Completion time did not differ among interventions.

This is evidence for a technique, not a general conversational-chart mandate. The task covered knowledge and comprehension in one educational setting, used online participants with limited domain context, and did not measure newsroom, BI, decision, accessibility, or long-term field outcomes. Its useful lesson is specific: an assistant that structures attention and asks the reader to reason can produce a different outcome from one that merely supplies an answer.

AI-generated misleading charts can measurably damage comprehension

A 2026 controlled experiment tested 48 readers on chart questions in two phases. Both groups first read correct charts. In the second phase, the control group continued to see correct charts while the experimental group saw misleading bar and line charts produced through an automated AI attack framework. Accuracy was 88.3% in the control condition and 71.9% in the misleading-chart condition. After adjustment for baseline chart-reading ability, education, chart type, and misleading technique, the odds of a correct answer were 0.266 in the misleading-chart condition.

This directly measures one kind of reader harm: a data-consistent chart whose design induces wrong answers. It does not estimate how common such charts are, whether ordinary users or systems create them accidentally, or how they affect trust and consequential decisions in news, BI, health, or social media. The study used horizontal and vertical bar charts and line charts in a controlled question-answering task. Its useful warning is narrower and stronger than a general claim about “AI misinformation”: visual design can remain faithful to the underlying table while materially reducing reader accuracy.

Accessibility is a delivered interaction, not a checkbox

Two studies outside AI chart generation provide an important adjacent baseline for evaluating delivered artifacts. In a controlled smartphone study, 26 low-vision participants completed bar- and line-chart tasks under five conditions. The full interactive treatment—which combined space compaction, personalization, and selective viewing—had 100% task completion, compared with 61.5% using the baseline screen magnifier. On a simple bar comparison, mean time fell from 531 seconds with the magnifier to 235 seconds with the full treatment. The study assured accurate chart extraction and used experimenter-provided phones, so it does not establish performance on arbitrary AI-generated charts or readers’ own devices.

A 2026 study then compared screen-reader text, audio-tactile exploration, and a refreshable tactile display with 10 blind adults completing 360 task episodes. Device-level accuracy did not differ significantly, but completion time and workload did. Chart type mattered even more: accuracy was 88.9% for pie charts, 81.1% for bars, 63.3% for lines, and 44.4% for scatterplots across the studied systems. No single modality was best for every task.

Neither study tests AI assistance. Together they replace an empty accessibility row with measured acceptance criteria: test the actual device, representation, interaction, chart type, task, time, error, and workload. A generated text description, responsive layout, or nominally accessible export is not itself a reader outcome.

Readers judge the data-generating claim too

In a 117-entry public discussion of a chart about AI-generated web content, readers quickly challenged the reliability of the AI-content detector behind the data. The reaction is a case, not a representative study, but it captures a key reader behavior: people do not only inspect axes and colors. They ask whether the measurement method can support the headline.

AI-assisted visualization therefore inherits the full author-reader contract: source quality, transformation choices, uncertainty, visual encoding, intended claim, and publication context. “The chart is accurate” is too small a promise.

Different contexts require different forms of assistance

The same assistant behavior can help in one environment and damage another.

Context Primary human purpose Useful assistance Particularly dangerous shortcut
Journalism and public explanation lead readers through evidence toward a bounded account exploration before publication, annotation variants, accessible implementation, source and data checks letting fluent generation substitute for reporting, authorial judgment, or audience testing
Operational dashboard maintain awareness and support timely response governed measures, anomaly explanation, stable layout edits, reader questions tied to source visuals changing metric definitions or visual hierarchy without operational ownership and regression checks
Analytical BI compare evidence and make decisions semantic-layer grounding, query generation, visible intermediate data, reusable verification confident answers over weak metadata, ambiguous measures, or hidden filters
Open-ended exploration form and test hypotheses many cheap views, branching, undo, direct manipulation, alternative transformations prematurely turning a plausible pattern into a narrative conclusion
Scientific visualization inspect domain-specific structures and support reproducible claims restricted representations, executable transformations, linked views, domain checks general visual plausibility standing in for specialized scientific correctness
Education and explorable explanation help a reader build intuition adjustable assumptions, guided interaction, multiple representations, feedback generating interaction without measuring what learners actually understand
Presentation and one-off communication communicate a deliberate argument in a constrained setting rapid drafts, formatting, accessibility, export, speaker-supporting variants producing generic polish without an intellectual trace of why this chart serves this audience

A useful common spine can preserve the question, data, decisions, artifact, provenance, and acceptance evidence. It should then delegate authoring and evaluation to the context. There is no reason to force a newsroom narrative, an operational alert, and an exploratory notebook through one model of autonomy or one acceptance score.

The same lifecycle exposes different evidence holes

The common spine becomes useful when it makes unlike contexts comparable without imposing one acceptance threshold. The table below applies the same sequence—question, data authority, generation, checking, correction, delivery, reader use, and sustainment—to four environments where the evidence is strongest or the consequences are clearest.

Environment What evidence exists now Where the lifecycle still disappears Acceptance evidence that belongs to this context
Journalism and public explanation Two current build diaries expose source work, task decomposition, data and geographic defects, correction, mobile checks, publishing, and proposed updates. Independent verification, editorial review, reader comprehension, accessibility, later refreshes, and civic use were not observed. Source-to-claim trace, editorial acceptance, delivered desktop/mobile/access states, audience comprehension, correction policy, and update ownership.
Governed BI and executive use Product contracts expose semantic models, queries, permissions, and review surfaces. Public testimony and analytics essays explain why narrow metrics and visible assumptions matter. Organization-owned definitions, routine task success, review queues, refusal, one-number incidents, trust recovery, and total cost are not independently measured together. Approved metric contract, exact query and filters, permission result, human owner, refusal behavior, incident escalation, and decision outcome.
Learning and explorable explanation A 117-person experiment measured immediate comprehension from proactive scaffolding; the classroom case above shows direct instruction rescuing a failed AI-assisted construction task. Delayed construction, unfamiliar transfer, learning after ordinary assistant use, and a complete learner-to-reader project remain sparse. Immediate and delayed unassisted performance, diagnosis, unfamiliar transfer, confidence calibration, and comprehension by the learner’s eventual audience.
Accessible and small-screen reading Controlled non-AI studies measure low-vision smartphone and blind nonvisual task outcomes. The accessibility sources above expose generated-description failure and organizational scale. No captured study joins AI-assisted authoring to disabled readers using their own devices and then measures verification, recovery, or harm. Co-designed alternatives, real devices and assistive technology, task success, errors, time, workload, independent verification, feedback, and remediation.

What current products reveal about the direction of travel

Product documentation cannot establish quality, but it shows where vendors are placing control and context:

This is a meaningful convergence: natural language is becoming one control surface inside a structured environment, not the whole environment. As baseline models improve, generic “make a good chart” instructions are likely to lose relative value. Maintained business semantics, provenance, deterministic tests, edit locality, delivery evidence, and reader-specific acceptance are less likely to be absorbed by the model because they belong to the situation, not the pretrained baseline.

A practical way to evaluate a tool or workflow

Do not begin with “Which model made the prettiest chart?” Begin with a real job and measure the entire path.

  1. Name the environment and audience. State whether this is exploration, governed BI, operational monitoring, public explanation, science, or a reusable application. Name who must read or maintain the result.
  2. Choose representative work. Use self-owned data and include awkward semantics, missing values, multi-step transformations, a correction request, and the actual delivery surface. Keep a simple task as a control.
  3. Record the baseline. Measure current human time, error rate, correction path, and delivery effort. A tool cannot be said to save time when the counterfactual is unknown.
  4. Separate first draft from acceptance. Record time to first useful candidate and time to accepted artifact, including prompt writing, waiting, cleanup, verification, and publishing.
  5. Inspect consequential choices. Can the person see the fields, filters, aggregations, transformations, code or semantic query, and generated interaction state? Can they make a precise local edit without regeneration?
  6. Test the delivered artifact. Recompute values, exercise controls, check accessibility and responsive behavior, and verify that exported or embedded output preserves the intended state.
  7. Test readers separately. Ask defined readers to state the main claim, supporting evidence, uncertainty, and next action. Measure correctness, time, confidence, and harmful misreadings. Do not use a model grader as the reader sample.

What remains missing—and what the new evidence partly fills

The gap list is no longer uniformly empty. New controlled studies now quantify one correction setting, one form of AI-generated chart harm, and several mobile and nonvisual reader outcomes. A production-workflow interview study adds a high-stakes dissemination boundary. None closes the central distance between a controlled task and routine delivered work.

Question What can now be said What remains open Highest-value next study
Can AI critique support work over time on self-owned material? Visualizationary followed 13 designers using self-selected data and tools over a three-to-five-day window and found average expert-rated improvement. observed work was roughly 90–150 minutes per participant, with no baseline, production delivery, or later maintenance multiweek within-person study on real commissions from intake through publication and update
What does correction cost, and when do people abandon? in 108 forced-error task episodes, 7 were not completed and participants prematurely declared completion 31 times; the novice study measured failed repairs; Visualizationary recorded 68 feedback uses and selective rejection the controlled tasks were engineered to fail and stopped at 15 minutes; there is still no field distribution of correction time, regeneration, rollback, abandonment, or downstream harm instrument every prompt, edit, wait, undo, verification, and abandonment against a direct-work baseline
Do local semantics improve correctness? verification research and current product contracts show why fields, measures, intermediate data, code, and verified queries matter; current BI cases report better trust with narrow curated metrics no independent head-to-head field test of a general assistant versus the same model with organization-owned semantics blinded task set with known local definitions, realistic ambiguity, and exact semantic-error scoring
Can the result be delivered and maintained? 17 biomedical-visualization practitioners described using AI chiefly for auxiliary tasks and keeping substantial generated imagery out of final scientific work; two 2026 newsroom diaries expose publishing, mobile checks, architecture, proposed updates, and unfinished maintenance; one direct second-maintainer episode exposes a creator-local refresh job and its replacement with a proper pipeline the interview study measured judgment rather than delivery; the diaries are self-reported prototypes; the handoff account is one unverified episode; controlled studies still stop before authenticated delivery, comparative maintenance, regression, or a later data refresh require real publishing, second-person handoff, one dependency change, a scheduled update, and an independently applied acceptance contract
What is the total cost? a 108-episode study measured task time and found no significant timing difference among conversational, stepwise, and phasewise interfaces; one 45-minute randomized exercise measured completion; the newsroom diaries add build durations, repeated correction, usage-limit waits, model-cost sensitivity, validation, and remaining work no common ledger joins human time, model inference or subscription cost, verification, failed generations, later delivery, maintenance, opportunity cost, and abandoned work against a current direct-work baseline prospective cost diary using time to first candidate and time to accepted maintainable artifact
Who benefits, by expertise and work context? the human-skills review separates consumption, construction, critique, and connection from data, domain, tool, and delivery resources; the learning chapter adds a positive immediate post-removal comprehension result for proactive scaffolding, adjacent randomized evidence that assisted performance can outrun learning, and ordinary visualization-retention evidence showing faster procedural decay; bounded studies show experienced practitioners filtering critique and using constrained implementation effectively “novice” remains inconsistently defined; expertise cells are small; delayed construction and far transfer remain mostly unmeasured; no captured longitudinal study causally estimates AI-driven visualization atrophy preregistered capability-profiled field trial crossing answer-oriented and metacognitive assistance with delayed unassisted transfer in multiple job environments
Does assistance help a reader understand—or harm understanding? the 117-person educational experiment found proactive scaffolding outperformed passive Q&A and a data story; a 48-person experiment found lower accuracy with AI-generated misleading charts than with correct charts both are controlled question-answering tasks; no independent decision quality, public communication, operational action, calibrated trust, or durable field learning independent reader trials in journalism, BI, education, and public services using delivered artifacts and consequential decisions
Does it work for disabled readers and on small screens? a 26-person low-vision smartphone study measured large differences across delivery treatments; a 10-person blind-reader study found modality and chart-type trade-offs across 360 tasks; a 12-person AI-assisted learning study found strong preference and spatial-model benefit for tactile + text + chat, but no measured accuracy lift; a 2025 position paper demonstrates information loss in model-generated description chains the AI-assisted study was short, small, and did not evaluate delivered artifacts or isolate a chart-vision specialist; broader disability groups, readers’ own devices, screen-reader and keyboard behavior, responsive reflow, verification, and remediation remain open co-designed AI-versus-non-AI delivery study on readers’ own devices with task success, errors, time, workload, independent verification, recovery, and qualitative experience
What does the work feel like in context—and what are people trying to protect? selected first-person accounts now distinguish mandate pressure, first-candidate momentum, repair interruption, reputational and scientific accountability, accessibility, craft, identity, and post-delivery disappointment; underheard-role evidence now includes ordinary spreadsheet allocation, provisional recipients, blind and low-vision learners, and one second maintainer consequential recipients, disabled readers on their own devices, non-English workplaces, displaced workers, and multiple second maintainers across contexts remain too faint; public accounts cannot estimate prevalence purposive multi-environment interviews and prospective diaries tied to actual artifacts, usage, corrections, decisions, handoffs, and abandonment
Which scaffolding still helps as models improve? proactive questioning beat passive answering in one reader experiment; deterministic perceptual checks and structured interfaces supported bounded authoring; skill projects publish provider-run tests few independent with-and-without ablations use the same current model, task, and harness, or repeat across model generations recurring benchmark and field panel that removes one scaffold at a time and records conflicts as well as gains

Semantic provenance, reader trust calibration, consequential decisions, and production regression remain especially thin. Harmful misinterpretation now has one controlled result, but no field estimate. The key experimental unit is not the generated chart. It is the complete creator-to-artifact-to-reader episode, with the job, environment, audience, and stakes named in advance.

Evidence guide: what the cited work actually is

Evidence Plain-language description What it supports here Important limit
2024 State of the Data Visualization Industry Online practitioner survey: 980 started, 763 completed, 825 answered the AI-use question; public respondent-level files permit role, task, and incumbent-tool cuts observed entry points for AI in visualization work and directional differences among this sample self-selected community survey with item-specific missingness; not representative adoption or product market share
2025 State of Analytics Engineering dbt survey of 459 practitioners and leaders on development, documentation, and natural-language data use general versus specialized assistant use and the gap between analytics creation and conversational consumption vendor-community sample, not visualization-specific, and based on self-report
DWP Microsoft 365 Copilot trial evaluation (2026 report) Official mixed-methods workplace evaluation with 1,716 licensed-user survey responses, 2,535 comparison responses, and 19 quota-sampled in-depth interviews routine work allocation, quiet non-use, Excel-specific limits, conditional trust, habit, time pressure, and stakeholder ownership nonrandom licence allocation, no pre-trial baseline, self-reported outcomes, and visualization as a subset of a broad office-assistant evaluation
Vibe Visualizing (2026 preprint) Think-aloud study of 20 visualization novices completing 60 ChatGPT chart tasks; 175 charts analyzed; initial prompts replayed on three current multimodal models novice prompting, design failures, confidence, verification, repair, and changing model behavior small controlled datasets; cross-model replay was not a live user study
How Good Is ChatGPT in Giving Advice on Your Visualization Design? (TOCHI 2025) Rating study plus 12-practitioner comparison of ChatGPT and human visualization advice brainstorming value, generic advice, context, expert preference practitioner sessions used 2023-era GPT-3.5
How Do Analysts Understand and Verify AI-Assisted Data Analyses? (CHI 2024) Design-probe study of 22 professional analysts and 52 verification workflows how people use code, data, explanations, and visual summaries to verify prepared tasks at one company; not full chart authoring
Randomized public-health analysis exercise (2025) 30 analyzed participants assigned to integrated ChatGPT or R/Stata plus ChatGPT for simulated epidemiological work speed-quality separation, tool integration, serious errors small, underpowered, simulated 45-minute exercise; no non-AI control
Data Formulator 2 (CHI 2025) Mixed-initiative chart-authoring prototype studied with eight corporate participants direct controls, visible transformations, history, inspection behavior reproduction tasks, GPT-3.5, no baseline or long-term use
DashChat (2025 preprint) Dashboard-prototyping system with formative interviews and a 28-person evaluation pre-data prototypes, stakeholder alignment, structured edits mockups rather than analytical correctness or production delivery
Visualizationary (2024 preprint) Thirteen designers iterated self-selected visualizations with LLM and deterministic perceptual feedback over a three-to-five-day window; three experts rated change longitudinal critique use, expertise differences, version tracking, and selective rejection of advice small, short study using GPT-3.5; no critique baseline, production delivery, or later maintenance
GenAI agents and visual-analytics comprehension (2024 preprint) Randomized 117-person comparison of a data story, passive Q&A, and proactive scaffolded dialogue with pre-, intervention-, and post-tests direct reader-comprehension evidence and the difference between answering and guiding controlled educational task, limited domain context and comprehension levels, no field deployment
Comparing LLM conversational and graphical interfaces for industrial decision tasks (2026 preprint) Twenty participants used a dashboard and chatbot across four simulated tasks, followed by questionnaires and semi-structured interviews conversational compression, dashboard overview and auditability, provisional reliance, and the separation between speed and confidence exploratory proxy sample dominated by computer-science students, short low-stakes tasks, one implementation of each interface, and human-LLM-assisted qualitative coding
AI-supported end-user development for visualization (2025) Eight interviews in one company followed by three profile-selected, 30-minute design-probe sessions basic, intermediate, and advanced users’ different prompting, transparency, and control needs exploratory single-company case; three sessions cannot establish usability or efficacy
How Do LLMs See Charts? (2026 preprint) Comparison of 24 people and three multimodal models on 60 synthetic charts why model and human chart reading are not equivalent synthetic data and narrow designer-intent labels
Trustworthy by Design (2025 preprint) 37 participants ranking charts and explaining perceived trust clarity, source, familiarity, integrity, and aesthetics as trust cues self-reported trust, not comprehension, correctness, or behavior
Vis Ex Machina (CHI 2021) Preregistered experiment with 114 people choosing identical recommendations labeled human or algorithmic contextual effects of provenance labels predates generative charts; study charts were researcher-authored
Improving Steering and Verification in AI-Assisted Data Analysis (UIST 2024) Eighteen experienced analysts completed 108 deliberately error-prone tasks across conversational, stepwise, and phasewise interfaces after a 15-person formative study premature acceptance, non-completion, correction controls, task time, and the cost of structured intervention forced-error 15-minute tasks using GPT-4 Turbo; no direct-work baseline, publication, or field abandonment
ChartAttack (2026 preprint) Controlled 48-person comparison of correct and AI-generated misleading bar and line charts, with baseline chart-reading phase and adjusted analysis direct reader-comprehension harm from selected visual misleaders controlled chart QA and limited chart types; not prevalence, calibrated trust, or consequential decisions
GenAI in biomedical visualization (TVCG 2026) Ninety-minute workflow interviews with 17 biomedical-visualization designers and developers spanning research, production, and dissemination where high-stakes practitioners use, restrict, or reject AI in real production workflows purposive Euro-North American qualitative sample; no observed delivery or audience outcomes
Low-vision charts on smartphones (IEEE VIS 2024) Fourteen formative interviews plus a controlled 26-person comparison of five ways to use bar and line charts on a smartphone mobile completion, time, error, usability, and workload as delivered-reader outcomes not an AI-generation study; accurate extraction was assured and participants used an experimenter-provided phone
Sound, Touch, or the Full Monty? (TACCESS 2026) Ten blind adults completed 360 tasks using screen-reader text, audio-tactile interaction, and a refreshable tactile display across four chart types modality, device, chart-type, time, accuracy, and workload trade-offs in nonvisual reading not an AI-assistance study; small controlled sample with limited training
Touching or Chatting (2026 preprint) Counterbalanced study of 12 blind or low-vision participants learning violin plots and clustered heatmaps with tactile + text + GPT-5.2 versus text + GPT-5.2; 263 substantive queries direct AI-assisted accessibility experience, complementary spatial and conversational roles, preference, and the separation between mental model and accuracy short small English-speaking US sample; no measured accuracy lift, long-term learning, or delivered-artifact evaluation
Second-maintainer dashboard discussion (2026 public testimony) One direct handoff episode plus 91 exposed replies about an AI-built dashboard failing during its creator’s absence hidden execution host, refresh lineage, semantic documentation, dependencies, ownership, and maintenance expectations self-selected and unverified testimony; no prevalence, comparison, or causal estimate
Public practitioner and reader threads Unrecruited discussions about current BI and spreadsheet work and chart reactions, including a 190-entry Excel discussion and a July 2026 dashboard-versus-chat discussion concrete jobs, delights, friction, abandonment, changes in peer learning, trust repairs, and reader questions that occur in practice self-selected testimony; establishes possibility of an experience, never prevalence or a population trend
Selected first-person creator logs and practitioner interviews Named accounts spanning early experimentation, guided learning, personal dashboards, visualization coursework, analytics practice, data journalism, accessibility, and analog craft how anticipation, delight, frustration, responsibility, refusal, and professional identity change across stages of work editorially selected and unusually articulate accounts; self-reported episodes are not comparative tests, population estimates, or independent artifact audits
Two 2026 Reuters Institute build diaries First-person accounts of rebuilding two data-journalism projects and constructing one national health dashboard with current coding agents production sequence, time compression, visible defects, correction, architecture, mobile checks, validation, ownership, and unfinished work three projects built and checked by their authors; no equal-budget baseline, independent audit, reader outcome, or maintenance follow-through
Playing Telephone with Generative Models (2025 position paper) One source visualization passed through four GPT-4o and Claude Sonnet 4 description-and-image chains mechanisms of information loss, fabrication, misplaced emphasis, poor design comprehension, under-description, overconfidence, verification disability, and compelled reliance authors explicitly frame the work as a provocation rather than a controlled experiment; no disabled-reader study or error-rate estimate
Doom or Deliciousness (CGF 2023) Semi-structured interviews with 21 experts in visualization or HCI, art or art history, and machine learning early field expectations about help and harm across datafication, transformation, visualization, and interaction convenience sample, mostly limited direct generative-tool experience, image-generation emphasis, and speculative framing
Dated field essays, interviews, and teaching or design cases Role-diverse assessments from visualization researchers and educators, analytics and library practitioners, a blind journalist, a creative studio, and data journalists purposes, workflow mechanisms, productive disagreements, lived constraints, and questions the controlled literature should test perspective and self-report are not prevalence, comparative efficacy, or reader-outcome evidence; positionality and publication date remain attached
Current product and project documentation Provider or maintainer descriptions for general assistants, spreadsheets, notebooks, editors, governed BI, data clouds, open-source systems, and skills the feature, data-access, target-user, control, and delivery contracts available in August 2026 self-reported feature scope and positioning, not independent adoption or comparative performance

Scope and limits

This review is a moment-in-time synthesis, not a market-share study or product ranking. Twenty-one human studies and structured workplace evaluations were read in full, along with practitioner surveys, current provider and open-source documentation, public practitioner and reader discussions, first-person production cases, and dated essays and interviews from several professional positions. Study populations, tasks, models, and evidence vintages remain visible because results are not directly interchangeable.

Public threads were used to discover and illustrate experiences that controlled research often omits. They were not sampled to estimate prevalence, and quoted claims were not independently reproduced. Vendor documents establish what a product says it supports, not that the feature is accurate or useful. Several important human-computer-interaction studies use 2021–2024 systems; their raw model comparisons are dated, while their observations about context, verification, and human control remain relevant hypotheses for current tools.

The strongest counter-reading is that better 2026 models may erase many errors reported in earlier work. The fresh novice study partly supports that: replayed prompts produced fewer design flaws with newer systems. It also found new failure modes from richer outputs, including render failures, misleading interpretations, latency, unclear interaction state, and greater verification burden. Capability is moving; the location of the human work is moving with it.

Update log