# The practitioner and reader experience of AI-assisted data visualization

Status: research snapshot, evidence cut 2026-08-15. Recheck by 2026-11-15, or
earlier after a material change in the major general-purpose, visualization,
or business-intelligence assistants discussed below.

Companions: [The state of AI-assisted data visualization research](/reports/research-review/),
[Human skills and banked gains in AI-assisted data visualization](/reports/human-skills-and-gains/),
and [Learning data visualization when AI can make the chart](/reports/learning-with-ai/).

This report asks a different question from the research review. Instead of
starting with architectures and benchmarks, it asks what it is like to use AI
to make or read a data visualization in August 2026. It covers the jobs people
are trying to perform, the kinds of tools available, measured studies of human
behavior, and public accounts of what feels remarkable or frustrating.

The short answer is that AI is already good at getting someone from a blank
page to a plausible first thing. It is much less reliable at carrying the work
through local meaning, correction, delivery, and audience understanding.
That distinction explains why a demo can be astonishing, a practitioner can
save real time, and the finished chart can still be wrong or unhelpful.

## Executive summary

AI-assisted visualization has three different success thresholds:

1. **Possible:** a model or tool can produce a chart, dashboard, analysis, or
   interactive artifact under some conditions.
2. **Accomplished:** a practitioner can turn that output into correct,
   inspectable, editable, and deliverable work at an acceptable total cost.
3. **Experienced:** creators understand and can control the process, while
   readers understand the result, notice its limits, and place appropriate
   trust in it.

Most public demonstrations stop at the first threshold. The strongest research
systems reach into the second by exposing transformed data, code, semantic
models, edit history, or direct controls. Very little work reaches the third:
independent reader comprehension, decision quality, calibrated trust,
accessibility, and mobile use remain sparsely measured.

Across the evidence, five conclusions are reasonably stable:

- **The reliable win is compression, not delegation.** AI can accelerate first
  drafts, chart reproduction, data preparation, routine code, documentation,
  and prototype variations. Those gains are real even when a human retains the
  analytical claim and final design.
- **Speed and polish are unsafe quality signals.** In a randomized data-analysis
  exercise, the faster integrated-AI group did not score better and produced
  fewer serious-error-free submissions. In a 2026 novice study, confidence and
  satisfaction remained fairly high despite pervasive chart flaws, task
  noncompliance, and some incorrect interpretations.
- **Context and inspectability determine whether a draft becomes usable work.**
  Systems do better when they can see governed metric definitions, field
  meanings, intermediate tables, code, worksheet state, or a restricted visual
  specification. Practitioners verify through artifacts they already
  understand; a mysterious result is difficult to trust or repair.
- **Expertise changes the experience rather than simply increasing the gain.**
  Experts can reject nonsense, spot broken assumptions, and redirect a model.
  They also recognize when writing the prompt, waiting for a regeneration, or
  cleaning the output costs more than direct work. Novices gain access but may
  be least equipped to notice consequential errors.
- **A good creator experience and a good reader experience are different
  outcomes.** Fast generation, visual appeal, and creator satisfaction do not
  establish that readers understand the intended claim. Readers use clarity,
  familiarity, sourcing, and aesthetics as trust cues, but those cues do not
  establish correctness.

A sixth conclusion is becoming visible but is less settled: **adoption is
entering through existing work rather than replacing the visualization stack.**
In the best available practitioner survey, AI use clustered around preparation,
analysis, coding, and ideation, while Excel remained the most widely used
visualization tool. Current BI testimony similarly describes chat taking ad hoc
questions while recurring dashboards retain shared definitions. This is a
directional reading of self-selected evidence, not a market-share estimate.

The practical implication is not “use AI” or “avoid AI.” Use it where the work
is reversible, the output remains inspectable, and the human can recognize a
bad result. Demand stronger grounding, deterministic checks, and explicit
reader testing as the semantic stakes, delivery costs, or audience consequences
rise.

## The gap that organizes the evidence

The same output looks different at each threshold. A generated dashboard can
render and still fail as work; a correct chart can still fail its reader.

| Threshold | The question being answered | Evidence that counts | Common false substitute |
| --- | --- | --- | --- |
| **Possible** | Can the system produce the requested class of output at least sometimes? | executable code, a rendered chart, task completion, supported feature contract | a polished screenshot or vendor description |
| **Accomplished** | Can someone finish the actual job correctly and maintain the result? | correct data and transformations, inspectable decisions, correction cost, delivery, reuse, total time and cost | first-pass speed, nonblank rendering, or a benchmark score alone |
| **Experienced — creator** | Can the person understand, steer, correct, and appropriately trust the process? | observed behavior, successful repairs, calibrated confidence, workload, abandonment, learning | satisfaction or stated preference alone |
| **Experienced — reader** | Does the final audience understand the claim and its limits and make an appropriate judgment? | comprehension, recall, decision quality, trust calibration, accessibility, behavior on the delivered device | model chart-reading, aesthetic ratings, or author confidence |

This is not a maturity ladder on which every visualization must climb. A
throwaway sketch may only need to be possible. A board metric, public-health
chart, or published explanatory graphic needs a defensible path through all
four rows.

## Who is actually reaching for what

The evidence is better at showing *where AI has entered the workflow* than at
naming a winning product. The [2024 Data Visualization Society survey](https://www.datavisualizationsociety.org/soti-report-2024)
was self-selected but unusually useful: 980 practitioners started it, 763
completed it, and 825 answered the AI-use question. Of those 825, 305 (37.0%)
said they had used AI in visualization work during the prior year, up 13
percentage points from 2023; 492 (59.6%) said no and 28 (3.4%) were unsure.

Among the 305 AI users, multiple selections were allowed. The work was weighted
toward the beginning of the lifecycle:

| Job reported by AI users | Respondents | Share of 305 AI users |
| --- | ---: | ---: |
| Prepare or clean data | 184 | 60.3% |
| Analyze data | 134 | 43.9% |
| Ideate or storyboard | 119 | 39.0% |
| Produce visualizations | 92 | 30.2% |
| Support a visualization team | 27 | 8.9% |

This is not a story of visualization specialists abandoning incumbent tools.
Excel was used often or sometimes by almost identical shares of AI users and
nonusers in the public microdata (66.9% versus 66.7%), and Tableau use was also
similar (43.9% versus 47.4%). AI users were more likely to report Python (42.6%
versus 27.4%), D3 (30.5% versus 13.4%), Observable (16.1% versus 5.7%), Figma
(40.0% versus 28.5%), and Datawrapper (19.7% versus 10.2%). The defensible
interpretation is that early adopters in this sample often already crossed
code, design, and publishing environments. Co-use does not show which tool
caused adoption.

A second, non-visualization-specific survey sharpens the split. In the
[2025 State of Analytics Engineering](https://www.getdbt.com/resources/state-of-analytics-engineering-2025),
80% of 459 respondents reported day-to-day AI use, up from 30% one year earlier.
Seventy percent used AI for analytics development, mainly through general
assistants such as ChatGPT, Claude, and Gemini; roughly 25% used specialized AI
inside development tooling. Only 30% were using natural language to consume
data, while another 29% wanted to and 23% had experimented. The sample came from
a vendor community and is not visualization-specific, but it reinforces a
plausible sequence: code and documentation adoption precede trusted
conversational consumption.

### The audience-to-environment map

The table separates four kinds of evidence that are often blurred together:
surveyed behavior, public cases, provider-defined target users, and our inferred
fit from the work surface.

| Who | What they appear to reach for first | Job they are trying to do | Status of the evidence |
| --- | --- | --- | --- |
| **Visualization practitioners with an existing stack** | a general assistant beside Excel, Tableau, Python, R, D3, Figma, or a publishing tool | clean data, debug code, learn a technique, storyboard, draft labels, make a first view | **Observed survey pattern.** The DVS sample measures AI jobs and incumbent-tool co-use, not specific AI-product share. |
| **Analytics engineers and code-capable analysts** | ChatGPT, Claude, Gemini, Copilot-style code help, then project-aware notebook or development agents | write and explain SQL or Python, document models, debug pipelines, generate a chart inside an analysis | **Observed survey pattern.** The dbt sample is vendor-adjacent and broader than visualization. |
| **Spreadsheet-native analysts, operations teams, and finance users** | AI inside Excel or Sheets; newer AI spreadsheets when code or live connections outgrow a grid | ask about a range, create formulas and charts, preserve familiar cells, hand work to colleagues | **Installed-base inference plus provider target.** Excel is widely observed; adoption of its AI features was not measured here. |
| **BI authors and data teams** | AI inside Power BI, Tableau, Looker, Hex, Sigma, or a governed data platform | author reports, define or reuse measures, inspect queries, answer stakeholder follow-ups, govern access | **Provider target plus public cases.** Independent head-to-head field evidence remains absent. |
| **Business consumers and executives** | conversational layers such as Genie, Spotter, Ask Sigma, Qlik Answers, Amazon Q, or Oracle’s assistant | retrieve a metric, ask why it changed, get an ad hoc cut, avoid navigating a complex report | **Provider target with self-selected testimony.** Current cases say chat complements recurring dashboards and loses trust quickly without curated definitions. |
| **Designers, journalists, and explanatory communicators** | general assistants for ideation or code, then Figma, Datawrapper, Flourish, or a newsroom-owned production stack | explore, create variants, annotate, improve accessibility, implement a story without surrendering editorial judgment | **Survey subgroup and workflow inference.** Journalism had 24 AI users among 47 valid DVS responses, too small and self-selected for a population rate. |
| **Developers and data-app builders** | coding agents, generated applications, declarative chart specifications, open-source agent frameworks, and reusable skills | make bespoke interaction, integrate data and UI, automate repeatable production, deploy an application | **Project positioning and inferred fit.** Repository activity shows availability, not routine user success. |

The July 2026 [public BI discussion](https://www.reddit.com/comments/1usnj1y/)
helps explain the business-consumer row without estimating its size. Reported
deployments use chat for exploration, follow-up questions, and requests that
previously became analyst emails; recurring dashboards still provide shared
monitoring and durable definitions. Participants repeatedly tied trust to a
narrow mart, approved metrics, visible refusals, and an owner for each measure.
Several described slow or failed pilots, and one reported that a single wrong
executive number ended trust. These are contemporary cases, not prevalence
evidence.

## The jobs: visualization is a workflow, not a prompt

“Make a chart” hides work before, during, and after visual encoding. Separating
the jobs matters because the same tool can be strong at one and harmful at the
next.

| Phase | Job | Useful assistance | Human responsibility that does not disappear |
| --- | --- | --- | --- |
| **Define** | **Frame the question** | restate a request, propose hypotheses, identify missing information | decide what matters, tie the work to a decision, reject an incoherent request |
| **Define** | **Acquire and govern** | find likely tables, draft access requests, describe schemas | choose authoritative sources, respect permissions, establish ownership and allowed use |
| **Work the data** | **Prepare and transform** | clean types, join, reshape, write SQL or Python, create routine calculations | define measures, resolve ambiguous fields, inspect exclusions and missingness |
| **Work the data** | **Analyze** | calculate summaries, fit standard models, compare groups, expose anomalies | choose a valid method, interpret uncertainty, know which causal claim is unavailable |
| **Work the data** | **Explore** | generate many cheap views, suggest cuts, translate follow-ups into operations | recognize spurious patterns, investigate causes, know when to stop branching |
| **Make the artifact** | **Sketch or reproduce** | remove blank-page work, copy a reference, create dashboard mockups and variants | decide whether the reference belongs in this context and what must change |
| **Make the artifact** | **Choose and encode** | suggest chart forms, mappings, aggregations, annotations, and emphasis | make the analytical and rhetorical choice for this data, claim, and audience |
| **Make the artifact** | **Critique and refine** | flag common perceptual problems, apply repetitive settings, draft labels and accessibility text | resolve composition, tone, brand, direct-manipulation details, and conflicting advice |
| **Make the artifact** | **Verify and debug** | expose code and intermediate tables, generate tests, replay interactions, compare results | state expected behavior, choose decisive checks, judge local semantic correctness |
| **Deliver** | **Package and publish** | assemble a report or app, write documentation, export standard formats | accept the delivered artifact, satisfy privacy and review gates, test the actual surface |
| **Deliver** | **Explain and interrogate** | provide tooltips, summaries, guided questions, and grounded conversational follow-ups | preserve the source claim, communicate uncertainty, evaluate reader understanding |
| **Sustain** | **Maintain and update** | regenerate routine artifacts, flag drift, propose regression checks | own dependency changes, semantic revisions, handoff, correction, and retirement |

The public cases are most consistent about work adjacent to the final visual:
SQL help, calculations, documentation, mockups, tooltips, bulk changes, and
alternative ideas. Claims that a system “built the dashboard” often compress
substantial context: an existing screenshot, prepared source tables, a semantic
model, a known template, or extensive cleanup.

## Listening to the field: what this work actually feels like

The accounts below are selected, not sampled. They come from first-person
project logs, interviews, public workplace discussions, and one qualitative
study of biomedical-visualization practitioners. Public handles, roles, and
outcomes are self-described unless a named publication or study says otherwise.
The accounts demonstrate that an experience was reported; they do not estimate
how common it is or prove that the artifact worked as described.

One distinction makes the listening more useful: **attitude and behavior are
not the same dimension**. A person can be delighted and unwilling to ship,
anxious and still using the tool, skeptical but grateful for boilerplate help,
or enthusiastic while retaining direct inspection and repair. A 17-person
[biomedical-visualization interview study](https://arxiv.org/pdf/2507.14494)
observed five such modes, from enthusiastic adoption through skeptical
avoidance. That specialized typology is not imposed on every account here. It
is a reminder not to reduce the field to boosters and opponents.

### Before the work: invitation, pressure, and hesitation

- An analytics team lead described leadership demanding AI “thought leaders”
  even though an end-user chatbot was unreliable and a pivot table was faster.
  The frustration was not simply with the model. It came from having to
  [perform AI adoption before identifying a useful job](https://www.reddit.com/comments/1srqm82/).
- Visualization coder [Jisell Howe](https://www.jisellhowe.com/posts/chatgpt-first-impressions/)
  initially feared that instant code would remove the rewarding journey from
  idea to customized chart. A concrete project changed that view: wrong syntax
  remained, but faster search and troubleshooting restored creative momentum.
- A master's student with little Python experience began a guided practical
  [unsettled by programming](https://www.linkedin.com/pulse/my-experience-exploring-generative-ai-python-data-fidel-nwaefulu-syguf).
  AI made a bar chart quickly; an instructor then supplied the design knowledge
  needed to improve it. The felt gain was access and accomplishment, not
  independent mastery.
- Biomedical-visualization specialists included eager adopters, interested
  non-users, cautious limiters, and principled avoiders. Resistance sometimes
  reflected scientific risk, threatened livelihoods, or pride in cultivated
  craft rather than unfamiliarity with the technology.

### The first candidate: access, velocity, and the rush of possibility

- A [Japanese independent creator](https://note.com/ktcrs1107/n/n25b998d0f4b1?hl=en)
  spent two days revising a personal dashboard and joked that the model revised
  their professional dignity too. A loose four-line instruction worked because
  the current files, desired fixes, and URLs were available. The surprise was
  that organized material mattered more than a polished prompt.
- A technologist taking a visualization class described a week with an agent as
  [“plain fun”](https://sef.kloninger.com/posts/claude-dataviz/). It helped process
  a 500 GB dataset, build tests, produce an exploratory dashboard, debug
  operations, and draft a presentation. The delight came from uninterrupted
  momentum across a whole project, not from one chart appearing.
- One experienced BI analyst said AI compressed an unfamiliar API task from an
  estimated week to a day, then cautioned that prior proficiency probably made
  that leverage possible. Another novice reported a much smaller but still
  meaningful gain: [writing cleaning and validation scripts](https://www.reddit.com/comments/1srqm82/)
  they could not otherwise have produced.

### The working loop: babysitting, interruption, and selective handback

- In one Power BI discussion, generated greenfield structure arrived quickly,
  while stale filters, bookmark identifiers, whitespace churn, and slow
  edit-preview cycles made local refinement painful. A practitioner called the
  approach good for new reports and clumsy for small changes; an experienced
  engineer said [describing the change could cost more than doing it](https://www.reddit.com/comments/1uccihy/).
- Analytics practitioner and educator [Christina Stathopoulos](https://www.datacamp.com/podcast/how-next-gen-data-analytics-powers-your-ai-strategy)
  described a useful analytical “sidekick” that still needed babysitting. In
  one episode it generated a chart and then stated the opposite of what the
  chart showed. The feared outcome was not an abstract benchmark failure but
  looking careless in front of stakeholders.
- Journalist [Jaemark Tordecilla](https://reutersinstitute.politics.ox.ac.uk/news/i-vibe-coded-complex-data-visualisation-and-analysis-dashboard-heres-what-i-learned)
  saw a recognizable site appear in about an hour. Local geography, evidence
  notes, usage limits, mobile behavior, architecture, and fact-checking then
  consumed repeated rounds. The work became most demanding exactly where the
  first candidate looked closest to done.

### Acceptance: reputation, accountability, access, and the right to refuse

- A dashboard consultant watched an agent reconstruct roughly a week's work
  from the finished screenshot and source tables in about 25 minutes. The
  result prompted a blunt professional question: [what is the consultant now
  charging for](https://www.reddit.com/comments/1vhrhys/)—the wrong portion, the
  ability to find it, or the invisible work already embedded in the reference?
- Biomedical-visualization practitioners welcomed boilerplate, inspiration,
  translation, and nontechnical assets but drew firm boundaries around anatomy,
  patient education, final images, and explainability. Some did not want to
  automate apparently mindless rendering because that was also where flow,
  control, and creative joy lived.
- Blind journalist [Johny Cassidy](https://reutersinstitute.politics.ox.ac.uk/blind-news-audiences-are-being-left-behind-data-visualisation-revolution-heres-how-we-fix)
  described exclusion from charts with missing or useless descriptions, then
  the relief of colleagues making access a shared habit. The experience makes
  clear why automatically emitted alt text is not yet accessible delivery:
  readers also need useful alternatives, feedback, and accountable correction.
- When public readers encountered a dramatic chart about AI-generated content,
  some accepted its story while others challenged its classifier, denominator,
  symmetry, and implied substitution. The [discussion became an argument about
  AI itself](https://www.reddit.com/comments/1ogp67o/) before the chart's method
  was resolved. Readers see a claim, not the creator's prompts and repairs.

### After delivery: use, learning, identity, and whether the saved effort mattered

- In the consultant discussion, another BI practitioner recalled an urgent
  dashboard that a manager praised and distributed to six colleagues. Three
  months later the usage record showed one manager view. Faster production can
  create an unused artifact faster; creator pride, stakeholder praise, and
  reader use are different outcomes.
- A BI developer described using AI not to generate the final chart but to
  [document SQL and calculations, review junior work, and rehearse stakeholder
  questions](https://www.reddit.com/comments/1kl0v0k/). The durable value was
  work around the dashboard that made it easier to explain and maintain.
- Freelance data journalist [Emilia Ruzicka](https://source.opennews.org/articles/data-hand-analog-datavis-self-reflection/)
  deliberately collected and drew personal data by hand. The slow, imperfect
  practice produced attention, experimentation, and a different relationship
  with the data. It is a reminder that not every friction is waste.
- In the leadership-pressure discussion, one commenter said their data team
  first feared automating itself away and later became the maintainer of the
  definitions and systems everyone relied on. Another described a leader
  trusting AI blindly and pushing bad recommendations back onto the team.
  Automation can elevate context work—or transfer more defensive work to the
  people who already own it.

### Listening past the original creator

A second listening pass pursued people who were mostly absent above: routine
spreadsheet workers, quiet non-users, recipients of generated analysis, and
the people expected to live with the result. It changes the picture in ways a
collection of creator success stories cannot.

| Setting and voice | Reported episode | What it makes visible | Limit |
| --- | --- | --- | --- |
| **Public-sector staff using Microsoft 365 in ordinary work** | A [DWP mixed-methods evaluation](https://www.gov.uk/government/publications/an-evaluation-of-dwps-microsoft-copilot-365-trial/an-evaluation-of-dwps-microsoft-365-copilot-trial) found that people routed work according to expertise, time, trust, data sensitivity, and habit. Experienced or data-heavy work often stayed in established methods; Excel users specifically reported trouble with complex data, charts, and formatting. One person valued the assistant yet was too busy to remember to use it consistently. | Non-use can be local and quiet. A person may value the overall assistant while declining it for the spreadsheet, chart, or stakeholder-owned task in front of them. | 1,716 user survey responses, 2,535 comparison responses, and 19 interviews provide substantial workplace evidence, but licences were not randomized, no pre-trial baseline existed, outcomes were self-reported, and visualization is a subset. |
| **A supply-chain analyst and a spreadsheet-help community** | A self-described intermediate Excel user said company encouragement made AI the first stop for reports and reduced visits to a peer forum. The [190-entry discussion](https://www.reddit.com/comments/1suy21w/) ranged from useful Power Query help to invented functions, wrong references, damaged formulas, and decisions never to ask again. Several people worried about losing the learning that comes from solving another person’s problem in public. | Assistance changes the support ecology, not only task time. Instant private answers can lower access cost while thinning public explanation, correction, and incidental learning. | One self-selected thread establishes that these experiences and concerns were reported. It cannot establish prevalence or a forum decline caused by AI. |
| **People making simulated industrial decisions** | Twenty participants used both a dashboard and an LLM conversational interface. They experienced conversation as compressed retrieval and mental offloading, but often regarded the result as provisionally useful rather than defensible. When speed was the criterion, 15 chose the chatbot and 3 the dashboard; when confidence was the criterion, all 18 decisive responses chose the dashboard. Several wanted both: overview to discover the question, conversation to retrieve or assemble the answer. | Usefulness and reliance separate for recipients just as delight and acceptance separate for creators. Extra visual interaction can be valuable when it exposes the values and relationships needed to defend a decision. | The [exploratory study](https://arxiv.org/pdf/2605.31224) used four short simulated tasks; 18 of 20 participants had computer-science backgrounds and were proxies rather than experienced industrial decision makers. |
| **Program recipients and an observer of an executive sales dashboard** | A [social-impact dashboard case](https://www.jaydavisportfolio.com/home/gooey-project) reports that program managers, evaluators, and funders wanted overview and drill-down, definitions, neutral language, comparative context, and review support. In a separate [named public critique](https://www.linkedin.com/posts/prithwirajmukherjee_someone-was-presenting-an-ai-generated-dashboard-activity-7425449434358112256-zeuy), telling an executive to verify regenerated charts by checking raw data felt like transferring the product’s verification burden to the person seeking a quick decision. | A recipient does not simply need a shorter summary. The path from alert to explanation and from claim to evidence must be proportionate to the recipient’s actual job. | The design case omits its interview count and instruments; the critique is one observer’s account and does not establish that the output was wrong or how the intended executive reacted. |
| **Blind and low-vision people learning unfamiliar chart forms** | In a [12-participant study](https://arxiv.org/pdf/2607.23065), people learned violin plots and clustered heatmaps with either tactile charts plus text and GPT-5.2 or text and GPT-5.2. Eleven preferred the tactile combination; one said the better mode depended on chart complexity; none preferred text and chat alone. Participants described touch as supplying the spatial model and chat as supplying flexible clarification. | Accessibility is not one generated description. Complementary modalities can answer different needs, and an open-ended assistant cannot always detect misunderstanding or supply questions a novice does not know to ask. | Preference and reported spatial understanding improved, but measured chart-understanding accuracy did not. This was a short study with a small English-speaking US sample, not a long-term learning or delivered-work evaluation. |
| **A second maintainer inheriting an AI-built dashboard** | A [public discussion](https://www.reddit.com/r/dataengineering/comments/1uhidmb/vibe_coded_dashboard_failing_on_a_friday/) begins with a dashboard failing while its creator was away. The inheriting maintainer found that refresh and scheduling depended on the creator’s laptop and replaced the setup with a proper pipeline. Replies described missing source and metric semantics, hidden dependencies, absent comments and documentation, and an implicit service expectation nobody had owned. | The maintenance boundary begins before handoff. Execution host, refresh lineage, definitions, dependencies, documentation, ownership, and rollback must be inspectable while the creator is still present. | This is one self-selected, unverified public episode plus discussion. It establishes that the experience was reported, not how common it is or whether AI rather than ordinary software practice caused every defect. |

Quiet withdrawal is therefore a gradient: task-level refusal, habitual fallback,
feature abandonment, substitution of one model for another, and provisional use
without confident reliance are different behaviors. A retained licence or active
session can hide all five.

The direct second-maintainer gap is no longer empty, but it remains extremely
thin. One episode is enough to make the hidden execution host and ownership
boundary concrete; it is not enough to estimate frequency, compare AI-assisted
and direct work, or describe what happens across organizations and tool stacks.

### What these voices reveal that aggregate capability evidence does not

- **Delight is often momentum:** the pleasure comes from staying in motion
  across chores, tools, and formerly blocked work—not necessarily from a better
  final chart.
- **Frustration is often interruption:** assistance becomes costly when a
  fluent practitioner must stop direct work to specify, wait, inspect, and
  repair a small change.
- **Fear has distinct objects:** job loss, reputation, lost craft, weakened
  learning, inaccessible delivery, scientific harm, and vendor dependence are
  not interchangeable concerns.
- **Trust is personal:** practitioners picture the stakeholder who will spot an
  obvious mistake, the patient affected by a wrong image, the client whose work
  bears their name, or the reader who cannot independently test a description.
- **Refusal can be expertise:** declining to use AI for a final visual may be a
  reasoned response to consequence, accountability, or the value of doing the
  supposedly low-level work directly.
- **The afterlife remains faint, not empty:** one direct handoff failure now
  exposes hidden execution and ownership assumptions, but routine reader use,
  maintenance across releases, correction, and retirement remain largely
  unobserved.

## What people in the field say has changed

Benchmarks tell us whether a system can perform a defined task. The perspectives
below explain why that task matters, what surrounds it in practice, and where
people disagree about the purpose of assistance. They include first-person build
diaries, practitioner essays, interviews, design-studio reflections, a classroom
exercise, and two research papers. They establish that an experience or argument
exists. Except where a measured study is named, they do not establish frequency,
comparative performance, or a general causal effect.

Across these selected perspectives, implementation is getting cheaper. The
substantive disagreement is what the saved effort should buy.

| Tension | What one perspective contributes | What the counterweight contributes | What context should decide |
| --- | --- | --- | --- |
| **Construction speed versus verification debt** | Two 2026 newsroom diaries describe work that once took weeks appearing in days or hours. They also report incorrect totals, misplaced geographies, missed notes, repeated correction, usage limits, mobile checks, and new validation machinery. | Analytics practitioner [Benn Stancil](https://benn.substack.com/p/the-smol-analyst) argues that a plausible chart cannot validate the calculation behind it; local definitions, executable lineage, and review remain necessary. | Semantic stakes, independence of the check, cost of a wrong result, and whether the complete pipeline—not only the first render—is owned. |
| **Lower implementation barriers versus more valuable judgment** | [Enrico Bertini](https://filwd.substack.com/p/what-can-ai-do-for-data-visualization) maps opportunities across acquisition, wrangling, analysis, data tours, and creative exploration. [Alberto Cairo](https://www.storytellingwithdata.com/blog/the-art-of-insight-in-conversation-with-alberto-cairo) treats easier code as capacity extension. | Both make question quality, synthesis, interpretation, and final analytical choice more important rather than less. Cairo also distinguishes exploratory, explanatory, essayistic, artistic, and poetic visualization purposes. | Purpose, audience, consequence, the person's capability profile, and whether the resulting claim can be inspected and challenged. |
| **Frictionless production versus productive friction** | Removing syntax and formatting work can let a person reach a candidate before the question goes cold. | In one [classroom exercise](https://teaching.virginia.edu/resources/data-visualization-with-and-without-ai), students spent 25–30 minutes failing with AI before direct instruction supplied the missing chart structure. An [analog data-journalism reflection](https://source.opennews.org/articles/data-hand-analog-datavis-self-reflection/) describes slow, manual work as a source of attention, experimentation, and care. | Whether the effort is obsolete mechanics or practice in structure, skepticism, memory, and meaning; whether assistance will later be withdrawn. |
| **Automated access versus independent verification** | Generative systems can draft descriptions and alternative representations at a scale that newsrooms have struggled to support. | [Elavsky and Xiong Bearfield](https://arxiv.org/pdf/2508.12192) show how model-to-model chart descriptions can lose data, source, uncertainty, design choices, and authorial purpose. Blind journalist [Johny Cassidy](https://reutersinstitute.politics.ox.ac.uk/blind-news-audiences-are-being-left-behind-data-visualisation-revolution-heres-how-we-fix) describes access as an organizational, multimodal, and reader-feedback problem—not an alt-text checkbox. | Whether the reader can test the account, reach the underlying data or authorial description, choose another modality, report a failure, and obtain remediation. |
| **Static charts versus conversational or divergent representations** | [Richard Brath](https://richardbrath.wordpress.com/2024/09/07/why-visualize-when-ai-can-find-insights-in-data/) asks which visualizations remain useful when a model can extract and organize insights. [Elijah Meeks](https://www.linkedin.com/posts/elijah-meeks_datavisualization-activity-7465183779041591296-EhLl) proposes making audience, intent, suggestion, interpretation, and conversation first-class framework objects. | [Domestic Data Streamers](https://www.domesticstreamers.com/art-research/work/what-can-genai-do-for-data-viz/) argues that visualization remains a distinct language for comparison, human experience, and creative divergence that a textual answer does not replace. | Reader task, need for overview or a shared evidence surface, qualitative meaning, desired novelty, and whether interaction leaves an auditable trace. |

An early formal baseline helps distinguish current observation from hindsight.
In [*Doom or Deliciousness*](https://www.mcnutt.in/assets/doom-n-fruit.pdf),
21 visualization, HCI, art, and machine-learning experts interviewed in 2023
anticipated rapid prototyping, personalization, and creative assistance alongside
bias, unreliable output, inattention, lost agency, and unearned trust. The study
used a convenience sample, foregrounded image-generation systems, and asked
participants to think beyond the technology they had used. It is useful as a
“what aged well?” comparison, not a verdict on current systems.

### Two current build diaries show where the time moved

The two journalism cases are unusually detailed accounts of current production.
They are still self-reports by the people who built and checked their own work.
Keeping the gain beside the incurred work prevents “built in a week” from
becoming a quality claim.

| Lifecycle stage | [Ottaviani: two reconstructed projects](https://reutersinstitute.politics.ox.ac.uk/news/weeks-work-days-how-i-rebuilt-two-data-journalism-projects-ai) | [Tordecilla: national health dashboard](https://reutersinstitute.politics.ox.ac.uk/news/i-vibe-coded-complex-data-visualisation-and-analysis-dashboard-heres-what-i-learned) | What neither case establishes |
| --- | --- | --- | --- |
| **Baseline and first candidate** | One map was rebuilt in two days after taking roughly three weeks in 2012; a second portal was attempted in a much more autonomous pass. | A repository became a recognizable site in about an hour; the full dashboard, maps, charts, analysis, and checking workflow took one week. | An equal-budget comparison with the best current direct workflow or another practitioner. |
| **Defects and correction** | Incorrect cluster and chart totals and incomplete bilingual text required source comparison and several correction rounds. | Highly urbanized cities rendered in the sea or in broken shapes; source notes and exact evidence were sometimes missed; long runs took shortcuts. | A field distribution of error, correction time, regression, premature acceptance, or abandonment. |
| **Architecture and control** | The author decomposed the work into about 20 tasks, ordered foundational work first, and kept code in an open repository. | The author separated extraction, data, presentation, insights, and fact-checking so a source error could be repaired and replayed. | That decomposition or a model-built validator independently improves correctness across projects. |
| **Delivery and readers** | The projects were published as working prototypes; the author identifies attention, understanding, and civic use as the remaining bottleneck. | Mobile rendering was checked and stakeholders reacted to the tool, but no reader-comprehension or decision study was conducted. | Audience understanding, accessibility, consequential use, or sustained adoption. |
| **Sustainment** | Proposed next steps include record-by-record reconciliation and scheduled updates. | Refactoring and longer-term maintenance remained after the one-week build. | Later refreshes, dependency changes, second-person handoff, incident response, or retirement. |

## The tools: a showcase of the current approaches

There is no single AI-visualization product category. A useful comparison asks
five literal questions: What can the assistant see? What can it change? Which
intermediate state can the user inspect? How local is a correction? What form
is actually delivered? The examples below are a current landscape, not a
ranking. Product pages establish feature contracts and intended users, not
independent quality or adoption.

### General assistants and specialist file chat

| Tool | Who it is positioned for | What it makes or changes | Where control returns to the person |
| --- | --- | --- | --- |
| [ChatGPT data analysis](https://help.openai.com/en/articles/9213685-extracting-insights-with-chatgpt-data-analysis) | anyone bringing a bounded file or table into a conversation | cleaning, executed analysis, common static or interactive charts, explanations, downloadable results | inspect generated code and tables; supply local definitions and move production work elsewhere |
| [Claude Artifacts](https://support.anthropic.com/en/articles/9487310-what-are-artifacts-and-how-do-i-use-them) | people who want a shareable custom visual or small application | generated HTML or application code, interaction, explanatory interfaces | inspect and edit the artifact; assume responsibility for state, accessibility, hosting, and maintenance |
| [Julius](https://julius.ai/home/ai-powered-business-intelligence) | file- or connection-oriented analysts seeking a dedicated data chat | analysis, charts, statistical models, reports, and presentations | continue through prompts or exported artifacts; independent repair and correctness evidence is sparse |

### Spreadsheet-native assistance

| Tool | Who it is positioned for | What it makes or changes | Important contract boundary |
| --- | --- | --- | --- |
| [Copilot in Excel](https://support.microsoft.com/en-us/excel/get-direct-answers-to-your-data-analysis-questions) | Excel beginners through advanced analysts | Python-backed answers and optional static chart or table insertion; an advanced mode can create refreshable Python cells | the direct-answer mode does not modify the workbook; inserted charts are static and non-refreshable |
| [Gemini in Google Sheets](https://support.google.com/docs/answer/14356410?hl=en-AU) | people already working in a shared spreadsheet | summaries, formulas, prompted edits, and editable charts inserted with supporting data on a new tab | the generated chart links to the new tab, not to subsequent changes in the original dataset |
| [Bricks](https://docs.thebricks.com/getting-started) | teams wanting spreadsheets, dashboards, slides, and collaboration in one workspace | grid-based analysis, live dashboard boards, presentation material | combines surfaces to reduce handoff, but current evidence is provider documentation rather than comparative use |
| [Quadratic](https://docs.quadratichq.com/) | spreadsheet users who also need Python, SQL, JavaScript, or live database connections | AI-generated code and queries, direct grid edits, code cells returning results to the sheet | exposes code and schema inside a familiar grid; asks users to manage a more technical workbook |

### Notebooks and visible canvases

| Tool | Who it is positioned for | What it makes or changes | Important contract boundary |
| --- | --- | --- | --- |
| [Hex AI](https://learn.hex.tech/docs/getting-started/ai-overview) | technical notebook authors, data teams, and consumers of published data apps | SQL, Python, chart, pivot, and Markdown cells; generated apps; semantic models; consumer chat | roles are explicit: technical users audit code, consumers query curated apps, managers govern semantic context |
| [Observable Canvases AI](https://observablehq.com/documentation/canvases/ai) | collaborative analysts who want AI work beside their own | new tables, transformations, SQL, and charts placed on the current canvas | the assistant adds new versions rather than silently editing or deleting; its view is bounded by viewport and schema |

### Guided storytelling editor

| Tool | Who it is positioned for | What it makes or changes | Important contract boundary |
| --- | --- | --- | --- |
| [Flourish AI](https://flourish.studio/product/ai-assistant/) | beginners, visualization practitioners, and teams refining a chart in a known editor | native editor settings for styling, labels, sources, annotation, and accessibility | proposed changes remain reversible in the editor; the assistant does not edit source data or invent missing intent |

### Governed BI and data-cloud assistants

| Tool | Primary surface | Intended reach | Distinguishing control or dependency |
| --- | --- | --- | --- |
| [Power BI Copilot](https://learn.microsoft.com/en-us/power-bi/create-reports/copilot-reports-overview) | report authoring inside Power BI | BI authors creating or modifying report pages | grounded in the model and report state; quality depends on maintained measures, descriptions, permissions, and review |
| [Tableau Agent](https://help.tableau.com/current/online/en-us/web_author_einstein_faq.htm) | Tableau web authoring | analysts building views and calculations in an existing workbook | works through the worksheet and data source, retaining a conventional authoring surface for correction |
| [Looker Conversational Analytics](https://cloud.google.com/looker/docs/conversational-analytics-overview) | natural-language questions over LookML | business users querying governed fields and measures | administrators own glossaries, defaults, verified queries, and semantic-model quality |
| [ThoughtSpot Spotter](https://www.thoughtspot.com/product/agents/spotter) | search and conversational analytics | business teams asking everyday or strategic questions | provider positions data teams as maintainers of governed models rather than answer queues |
| [Ask Sigma](https://www.sigmacomputing.com/blog/announcing-ask-sigma) | conversational analysis over a cloud data workspace | non-specialists and analysts exploring organization data | shows data sources, formulas, filters, and analysis steps and allows step-level edits instead of full regeneration |
| [Qlik Answers](https://help.qlik.com/en-US/evaluation-guides/Content/ai/qlik-answers.htm) | structured and unstructured enterprise knowledge | business users asking questions and authors assembling analytic output | combines RAG, chart and dashboard agents, permissions, and audit logs; evidence remains provider-authored |
| [Databricks Genie](https://docs.databricks.com/aws/en/genie/talk-to-genie) | a curated, no-code data chat | business users asking questions beyond existing dashboards | returns SQL-backed answers and visualizations; authors monitor, review, and refine instructions and definitions |
| [Amazon Q in QuickSight](https://docs.aws.amazon.com/quick/latest/userguide/quicksight-gen-bi.html) | BI authoring, Q&A, executive summaries, and data stories | authors and consumers already in the AWS BI surface | integration reduces handoff; correctness still inherits datasets, topics, permissions, and author review |
| [Oracle Analytics AI Assistant](https://docs.oracle.com/en/cloud/saas/analytics/26r1/fawug/oracle-analytics-ai-assistant.html) | natural-language analytics in Oracle | authors who prepare metadata and consumers who request charts or narratives | explicitly separates the author who prepares data and metadata from the consumer who asks questions |

### Mixed-initiative and open-source systems

| Project | Approach | Why it matters | Evidence boundary |
| --- | --- | --- | --- |
| [Data Formulator](https://github.com/microsoft/data-formulator) | direct encoding controls plus natural-language transformation, branching, visible derived data, and report sharing | treats AI as one control surface in an inspectable analysis environment | research evaluations are small; the current repository is more capable than the studied GPT-3.5 system |
| [Lumen](https://github.com/holoviz/lumen) | declarative data pipelines, visualizations, dashboards, and specialist agents that can be serialized and edited | makes generated transformations and output inspectable and reusable across chat, notebooks, and dashboards | open-source feature surface, not independent evidence of routine field success |
| [Vizro](https://github.com/mckinsey/vizro) | low-code Python dashboard specification with code escape hatches and agent-facing tools | constrains generation to a production-oriented component system while preserving extension and deployment paths | a reusable framework reduces some errors; acceptance still requires tests and design judgment |
| [AntV chart visualization skills](https://github.com/antvis/chart-visualization-skills) | retrievable, version-specific instructions and declarative chart or code generators for several visualization libraries | shows how a harness can supply library syntax, chart vocabulary, and guardrails at generation time | the project reports its own 174-case testing; independent and current-model ablations are still needed |

Six techniques now cut across product classes: unrestricted code generation;
semantic grounding in governed fields and measures; structured or declarative
intermediate representations; direct manipulation and local undo; deterministic
render or data checks; and proactive reader scaffolding. These techniques can
coexist. The meaningful choice is which one contains the most expensive failure
for the actual job.

A practitioner may therefore ask a general model for dashboard concepts, use a
coding agent to generate measures, edit the layout manually in BI software, and
let a reader-facing assistant explain the final report. That is one workflow
with four evidence and control boundaries—not one assistant experience.

## What feels amazing to creators

### A first draft appears before the idea has cooled

The most consistent delight is immediacy. A person can move from a question or
rough picture to something concrete enough to react to. In the DashChat study,
participants valued rapid multi-chart mockups and simulated data as a way to
negotiate dashboard requirements before production data existed. Data
Formulator participants could combine direct encoding choices with short
transformation prompts and complete all 16 reproduction tasks, although the
eight-person study did not include a conventional-tool baseline.

This is valuable because a draft changes the conversation. “Show me three ways
to compare this” is easier to critique than an abstract discussion of a future
dashboard. The benefit is not that the first output is right; it is that the
cost of making a candidate has fallen.

### Tedious, legible work compresses well

Practitioners repeatedly describe value in SQL, calculations, documentation,
bulk renaming, standard chart setup, tooltips, formatting, and alternative
approaches. These tasks have inspectable outputs and a relatively short path to
correction. In one [public BI discussion](https://www.reddit.com/comments/1kl0v0k/),
contributors described using AI to explain dashboard queries in documentation,
prepare stakeholder questions, review junior work, critique mockups, and
accelerate unfamiliar API or Python tasks. This is public testimony, not a
survey, but it gives the broad phrase “AI in BI” concrete content.

### AI can bridge representations a person does not fluently command

The 22-analyst verification study found that people moved among natural-language
explanations, code, original data, intermediate data, results, and visual
summaries. When Python was unfamiliar, some participants translated the logic
into SQL or operations they knew. AI is useful here as a representation bridge:
it can express the same intent as prose, code, a table, or a visual candidate.
The person still needs at least one representation in which they can evaluate
the result.

### A bounded integrated tool can be genuinely faster

In a randomized public-health analysis exercise with 30 analyzed participants,
the integrated ChatGPT environment finished faster than the distributed
R/Stata-plus-ChatGPT workflow: a median 38 versus 45 minutes. The integrated
group’s visualizations were also judged more meaningful and correctly labeled.
That is real evidence for reducing tool-switching and syntax work in this
specific 45-minute exercise. It is not evidence that the integrated group’s
overall work was safer or better; that result points in the other direction.

### A critic can improve work without becoming the designer

[Visualizationary](https://arxiv.org/pdf/2409.13109) followed 13 designers—six
novice, four intermediate, and three expert—across a three-to-five-day window
as they iterated at least five versions of a visualization made with their own
data and preferred tool. Its GPT-3.5 critique was combined with deterministic
perceptual filters, a hierarchical report, and version tracking. Three external
experts later rated the final work as improved on average (3.69 on a five-point
scale where three meant no change).

The system did not supply a universally correct edit list. Participants used 68
pieces of feedback, ignored advice that conflicted with intent or seemed wrong,
and did not always end with their best-rated version. Intermediate and expert
designers generally converted critique into better revisions more effectively
than novices. The small, short study has no conventional-critique baseline, but
it demonstrates a credible role for AI as inspectable counsel inside an
existing practice rather than as an autonomous author.

## What makes the work hard

### The model can make an invisible analytical decision

Aggregation, filtering, grouping, baseline, and scale choices can be buried in
generated code or made implicitly from an ambiguous prompt. In the 2026 novice
study, 13 of 20 participants issued ambiguous or problematic prompts, and the
system sometimes chose inconsistent aggregations or supplied interpretations
the user could not verify. A chart can look conventional while encoding a
different question from the one its creator thinks it answers.

Governed BI environments try to narrow this gap through semantic models,
descriptions, verified queries, permissions, and authored instructions. The
current Looker documentation is unusually explicit: the model maps natural
language onto maintained LookML fields and lets administrators supply business
glossaries, defaults, and verified queries. That is an architecture contract,
not proof of perfect answers. It also makes the organizational dependency
visible: conversational analytics inherits the quality of the semantic model
someone maintains.

### Correction can be harder than creation

Natural language is a powerful coarse control and often a poor fine control.
In the novice study, 11 of 16 attempts to fix clutter and eight of nine attempts
to repair unusable charts failed. More capable systems generated richer
interfaces, but also introduced longer waits, unclear selected states, broken
controls, and cases where a model interpreted a blank rendering as if it
contained data.

Public Power BI accounts add a practical version of the same problem. In a
[35-entry thread](https://www.reddit.com/comments/1uccihy/), contributors
described useful greenfield structure alongside stale filters, bookmark IDs,
whitespace cleanup, slow preview cycles, and minor changes that were faster to
make manually. One contributor summarized the division as trusting and
verifying generated code while retaining human judgment over whether business
users would accept the report. These are cases, not frequencies.

Another practitioner reported a generated inline calculation that produced a
desired bar chart but could not be located or edited through familiar Power BI
mechanisms. The issue is not merely whether the number happened to be right.
The generated implementation broke the user’s ability to inspect, maintain,
and explain it.

The longitudinal Visualizationary study found a quieter form of the same cost:
feedback could be verbose, vague, or based on a mistaken reading of intent.
Novices in particular struggled to translate broad advice into a precise edit,
and the final version was not always the highest-rated one. A critique system
therefore needs triage, rationale, and edit locality; more feedback is not the
same as a better correction path.

A 2024 controlled study of AI-assisted data analysis gives this problem an
actual, if deliberately artificial, denominator. Eighteen experienced analysts
completed 108 task episodes designed so that the model would make two errors
unless the person found and repaired them. Seven episodes were not completed
within 15 minutes. In 31 episodes—28.7% of the study tasks—the participant said
the work was complete while an issue remained. Interfaces that exposed editable
assumptions and plans improved participants' sense of control, but the study did
not detect a difference among conditions in task success, completion time, or
verification hints.

This is evidence that premature acceptance and non-completion occur even among
experienced analysts who have been warned to look for errors. It is not a field
failure rate: every task was engineered to be troublesome, the researcher
supplied correctness checks and hints, and no episode continued through
publication or abandonment in ordinary work.

### Total cost includes waiting, prompting, cleanup, and model usage

First-pass generation time omits the rest of the loop. Practitioners mention
prompt construction, regeneration latency, application refreshes, token or
credit consumption, deployment cycles, and the cognitive cost of deciding
whether a strange result is a model error or their own misunderstanding. A
one-minute generation can be slower than a ten-second direct edit; a 25-minute
dashboard rebuild can be impressive and still depend on an existing screenshot,
prepared tables, and substantial layout review.

The useful unit is therefore total human-plus-system time to an accepted
artifact, including correction and verification—not seconds to first render.

The same 108-episode study illustrates why interface improvements should not be
translated automatically into time savings. Mean completion time was 543
seconds for the conversational baseline, 588 seconds for stepwise decomposition,
and 658 seconds for phasewise decomposition; the differences were not
statistically significant. More visible structure made correction feel more
controllable while also creating intervention and information costs. The study
measured human task time, not model inference, subscription cost, later cleanup,
or maintenance, so it still does not provide a total-cost ledger.

### Production adds failure surfaces that studies often remove

Controlled studies usually provide clean data, bounded tasks, and a working
prototype. Real work adds permissions, confidential data, incomplete semantics,
unreliable connectors, browser behavior, responsive layout, organizational
review, publishing systems, and future maintenance. The 2025 randomized
exercise recorded upload and browser problems even in a simulated setting.
Practitioners in public threads describe network policy, application refresh,
deployment, and cost constraints that are normally absent from benchmark
scores.

Interviews with 17 biomedical-visualization practitioners show where a
high-stakes production boundary is currently being drawn. Thirteen already used
generative AI somewhere in their workflow, primarily for auxiliary work such as
research, moodboards, boilerplate code, background assets, captions, or
translation. All opposed incorporating substantial AI-generated visual content
into final scientific work, citing validation, accuracy, copyright, provenance,
and accountability. That is evidence about real production and dissemination
judgment, not delivery efficacy: the study did not observe projects through
handoff, publication, reader use, refresh, or maintenance.

## Speed and quality separate in the measured studies

The most cautionary evidence is not that systems always fail. It is that their
success signals can disagree.

| Study | What improved or looked good | What did not follow |
| --- | --- | --- |
| **Randomized public-health analysis exercise, 2025** — 30 analyzed participants | Integrated AI users finished seven minutes faster at the median; their charts were judged more meaningful and correctly labeled. | Overall scores were not significantly different. Only 6.7% of integrated submissions versus 26% of distributed-tool submissions were free of serious errors. Small, underpowered study; no non-AI control. |
| **Vibe Visualizing, 2026** — 20 novices, 60 sessions, 175 charts | Participants could repeatedly produce charts and reported mean confidence of 3.73/5 and satisfaction of 3.93/5. Replayed prompts showed better visual output from newer models on some measures. | Every chart had at least one design flaw; 52 of 60 sessions had fatal task noncompliance; 22 charts were unusable; incorrect insights appeared in 12 sessions. Verification was rare and many repair attempts failed. |
| **Data Formulator 2, CHI 2025** — eight corporate participants | All participants completed 16 reproduction charts; direct controls, short prompts, visible transformed data, history, and multiple inspection artifacts supported work. | Six of eight still needed hints. The tasks were reproduction, not open analysis; no conventional-tool baseline or self-owned data; trust in simple steps could carry too easily into complex descendants. |
| **DashChat, 2025 preprint** — formative interviews plus 28-person evaluation | Rapid pre-data dashboard prototypes helped make requirements and layouts discussable; designers rated it easier and more effective than a lightly taught Tableau condition. | The comparison favored the familiar purpose-built workflow and measured mockups, not real-data correctness, deployment, or ongoing dashboard use. Precise edits and domain conventions remained difficult. |

These results should not be pooled into a universal score. They agree on a
more useful point: assistance can improve access, speed, or perceived ease
without proportionately improving semantic correctness, error freedom, or
delivery.

## Expertise is both leverage and burden

The simple story that novices benefit and experts do not is wrong. So is the
story that experts automatically get the best results.

Visualization expertise is not one rank. The strongest current literacy
framework separates the ability to **consume**, **construct**, **critique**, and
**connect** a visualization to its context. Real work also draws separately on
data and statistical knowledge, domain semantics, visual design, implementation,
situated judgment, and delivery experience. A person can be a strong business
reader and a weak programmer, a domain expert and a novice chart author, or an
expert implementer who does not know the local measure definitions. AI removes
different friction for each profile.

The [human-skills companion](/reports/human-skills-and-gains/)
therefore uses a stricter definition of a gain. Access, productivity, artifact
quality, independent learning, verification, reader outcome, and sustained work
are separate. Confidence or an AI-assisted final project does not establish
learning; time to a first render does not establish productivity to an accepted
artifact.

The learning companion follows that distinction across time. It finds that
implementation, translation, and on-demand explanation are becoming easier,
while delayed independent construction remains mostly unmeasured. It recommends
removing inspectable production friction while preserving prediction,
data-to-encoding reasoning, comparison, diagnosis, local repair, retrieval, and
unassisted transfer.

### Experienced practitioners can filter and redirect

In the comparison between ChatGPT advice and human visualization experts,
practitioners valued AI for rapid brainstorming, broad option lists, and a
neutral first pass. One experienced participant’s description was roughly: ask
for 20 ideas, discard 16, use four. That workflow is valuable because the
person can perform the filtering. In the analyst verification study, people
with data experience inspected distributions and values, while coders inspected
code and others translated logic into familiar operations.

The longitudinal designer study adds direct evidence. Intermediate and expert
participants were better able to select among critiques and enact useful
changes, while novices needed more help turning high-level feedback into design
operations. A 2025 single-company case study reached a compatible but much
weaker result: eight interviews produced basic, intermediate, and advanced
profiles, then one person from each profile used a prototype for about 30
minutes. The basic user needed initial prompting support and missed some
interaction affordances; the advanced user demanded source transparency,
encountered a wrong scatterplot, and formulated more intricate corrections.
Three probe sessions cannot establish efficacy, but they show why one interface
cannot equate “easier” with less visible state for every user.

### The same expertise exposes prompt overhead

Human experts were still preferred to ChatGPT for accuracy, helpfulness,
reliability, adaptability, context, and actionable advice in the 12-practitioner
study. The model produced broad, agreeable recommendations; experts narrowed
the scope, built common ground, and raised issues the practitioner had not
known to ask about. The model sessions were conducted in 2023, so their raw
quality comparison is dated. The contextual lesson remains visible in newer
public accounts: experienced people often reserve AI for repetitive or
unfamiliar work because describing a familiar small edit can cost more than
doing it.

### Novices gain access without gaining an error detector

The 2026 novice study is the clearest warning. Participants could produce
plausible charts, but verification intent was low, only three verification
attempts were observed, and satisfaction remained positive. Some users followed
inappropriate model suggestions. Poor visualization and data literacy affected
both prompting and the ability to diagnose the result.

This creates an expertise paradox: the people for whom the interface removes
the largest access barrier may be the least able to tell when it has quietly
answered the wrong question. Better defaults help, but they do not eliminate
the need to expose the decision the default made.

## The reader receives a claim, not a generation process

Creators experience prompts, waiting, code, and corrections. Readers usually
see only the final artifact in a report, article, dashboard, presentation, or
social post. They infer care and authority from what is visible.

### Trust is assembled from cues

In a 2025 preprint, 37 US participants ranked static charts from news, science,
government, and infographic contexts and explained their choices. Clarity was
mentioned by 31 of 37 participants; source citation, familiar chart forms,
integrity, and visual polish also mattered. Priorities were consistent within
many individuals and different across individuals. Infographics polarized
viewers: some saw accessibility, others saw promotion or clutter.

The study measured deliberative, self-reported trust—not truth detection,
comprehension, or behavior. That distinction is decisive for AI-generated
charts. A clean chart can signal professional care even when no one inspected
its transformations. A cited source can increase confidence without proving
that the chart represented it faithfully.

### An AI label does not have one predictable effect

The preregistered 2021 Vis Ex Machina experiment showed identical chart panels
with human or algorithm labels to 114 participants. Before the task, 60%
preferred human recommendations and 15% preferred algorithmic recommendations;
actual choices were approximately even. Data relevance dominated, but some
participants followed their prior source preference. An algorithmic label
suggested precision or error-freedom to some people and lack of human judgment
to others.

That study predates generative visualization and does not tell us whether to
label current charts “AI-generated.” It establishes that provenance cues
interact with prior beliefs and task context; disclosure is not a universal
trust switch.

### A model reader is not a human-reader test

In a 2026 study of 60 synthetic charts, three multimodal models matched a narrow
designer-intent label more often than 24 human participants. The models tended
to enumerate structure and values; people formed trend narratives and were
more affected by layout and overlap. The finding is not that models are better
chart readers. It is that the metric rewarded a task aligned with machine
decoding while human reading pursued a different form of meaning.

Using a model to inspect a chart can catch useful defects. It cannot establish
that a journalist’s audience understood the story, a dashboard user noticed an
alert, or a student learned the concept.

### Reader assistance can guide comprehension, not only answer questions

A randomized experiment with [117 higher-education participants](https://arxiv.org/pdf/2409.11645)
compared three ways to support interpretation of a bar chart, communication
network, and ward map: a conventional data story, a passive GenAI agent that
answered questions, and a proactive agent that asked educator-authored
scaffolding questions and gave feedback. All three groups improved from their
unsupported baseline. After the assistance was removed, the proactive group’s
median score was 6 of 6, compared with 5 for both the data-story and passive
groups; the between-group effect was statistically significant and medium in
the authors’ analysis. Completion time did not differ among interventions.

This is evidence for a technique, not a general conversational-chart mandate.
The task covered knowledge and comprehension in one educational setting, used
online participants with limited domain context, and did not measure newsroom,
BI, decision, accessibility, or long-term field outcomes. Its useful lesson is
specific: an assistant that structures attention and asks the reader to reason
can produce a different outcome from one that merely supplies an answer.

### AI-generated misleading charts can measurably damage comprehension

A 2026 controlled experiment tested 48 readers on chart questions in two phases.
Both groups first read correct charts. In the second phase, the control group
continued to see correct charts while the experimental group saw misleading
bar and line charts produced through an automated AI attack framework. Accuracy
was 88.3% in the control condition and 71.9% in the misleading-chart condition.
After adjustment for baseline chart-reading ability, education, chart type, and
misleading technique, the odds of a correct answer were 0.266 in the
misleading-chart condition.

This directly measures one kind of reader harm: a data-consistent chart whose
design induces wrong answers. It does not estimate how common such charts are,
whether ordinary users or systems create them accidentally, or how they affect
trust and consequential decisions in news, BI, health, or social media. The
study used horizontal and vertical bar charts and line charts in a controlled
question-answering task. Its useful warning is narrower and stronger than a
general claim about "AI misinformation": visual design can remain faithful to
the underlying table while materially reducing reader accuracy.

### Accessibility is a delivered interaction, not a checkbox

Two studies outside AI chart generation provide an important adjacent baseline
for evaluating delivered artifacts. In a controlled smartphone study, 26
low-vision participants completed bar- and line-chart tasks under five
conditions. The full interactive treatment—which combined space compaction,
personalization, and selective viewing—had 100% task completion, compared with
61.5% using the baseline screen magnifier. On a simple bar comparison, mean time
fell from 531 seconds with the magnifier to 235 seconds with the full treatment.
The study assured accurate chart extraction and used experimenter-provided
phones, so it does not establish performance on arbitrary AI-generated charts
or readers' own devices.

A 2026 study then compared screen-reader text, audio-tactile exploration, and a
refreshable tactile display with 10 blind adults completing 360 task episodes.
Device-level accuracy did not differ significantly, but completion time and
workload did. Chart type mattered even more: accuracy was 88.9% for pie charts,
81.1% for bars, 63.3% for lines, and 44.4% for scatterplots across the studied
systems. No single modality was best for every task.

Neither study tests AI assistance. Together they replace an empty accessibility
row with measured acceptance criteria: test the actual device, representation,
interaction, chart type, task, time, error, and workload. A generated text
description, responsive layout, or nominally accessible export is not itself a
reader outcome.

### Readers judge the data-generating claim too

In a [117-entry public discussion](https://www.reddit.com/comments/1ogp67o/)
of a chart about AI-generated web content, readers quickly challenged the
reliability of the AI-content detector behind the data. The reaction is a case,
not a representative study, but it captures a key reader behavior: people do
not only inspect axes and colors. They ask whether the measurement method can
support the headline.

AI-assisted visualization therefore inherits the full author-reader contract:
source quality, transformation choices, uncertainty, visual encoding, intended
claim, and publication context. “The chart is accurate” is too small a promise.

## Different contexts require different forms of assistance

The same assistant behavior can help in one environment and damage another.

| Context | Primary human purpose | Useful assistance | Particularly dangerous shortcut |
| --- | --- | --- | --- |
| **Journalism and public explanation** | lead readers through evidence toward a bounded account | exploration before publication, annotation variants, accessible implementation, source and data checks | letting fluent generation substitute for reporting, authorial judgment, or audience testing |
| **Operational dashboard** | maintain awareness and support timely response | governed measures, anomaly explanation, stable layout edits, reader questions tied to source visuals | changing metric definitions or visual hierarchy without operational ownership and regression checks |
| **Analytical BI** | compare evidence and make decisions | semantic-layer grounding, query generation, visible intermediate data, reusable verification | confident answers over weak metadata, ambiguous measures, or hidden filters |
| **Open-ended exploration** | form and test hypotheses | many cheap views, branching, undo, direct manipulation, alternative transformations | prematurely turning a plausible pattern into a narrative conclusion |
| **Scientific visualization** | inspect domain-specific structures and support reproducible claims | restricted representations, executable transformations, linked views, domain checks | general visual plausibility standing in for specialized scientific correctness |
| **Education and explorable explanation** | help a reader build intuition | adjustable assumptions, guided interaction, multiple representations, feedback | generating interaction without measuring what learners actually understand |
| **Presentation and one-off communication** | communicate a deliberate argument in a constrained setting | rapid drafts, formatting, accessibility, export, speaker-supporting variants | producing generic polish without an intellectual trace of why this chart serves this audience |

A useful common spine can preserve the question, data, decisions, artifact,
provenance, and acceptance evidence. It should then delegate authoring and
evaluation to the context. There is no reason to force a newsroom narrative,
an operational alert, and an exploratory notebook through one model of
autonomy or one acceptance score.

### The same lifecycle exposes different evidence holes

The common spine becomes useful when it makes unlike contexts comparable
without imposing one acceptance threshold. The table below applies the same
sequence—question, data authority, generation, checking, correction, delivery,
reader use, and sustainment—to four environments where the evidence is strongest
or the consequences are clearest.

| Environment | What evidence exists now | Where the lifecycle still disappears | Acceptance evidence that belongs to this context |
| --- | --- | --- | --- |
| **Journalism and public explanation** | Two current build diaries expose source work, task decomposition, data and geographic defects, correction, mobile checks, publishing, and proposed updates. | Independent verification, editorial review, reader comprehension, accessibility, later refreshes, and civic use were not observed. | Source-to-claim trace, editorial acceptance, delivered desktop/mobile/access states, audience comprehension, correction policy, and update ownership. |
| **Governed BI and executive use** | Product contracts expose semantic models, queries, permissions, and review surfaces. Public testimony and analytics essays explain why narrow metrics and visible assumptions matter. | Organization-owned definitions, routine task success, review queues, refusal, one-number incidents, trust recovery, and total cost are not independently measured together. | Approved metric contract, exact query and filters, permission result, human owner, refusal behavior, incident escalation, and decision outcome. |
| **Learning and explorable explanation** | A 117-person experiment measured immediate comprehension from proactive scaffolding; the classroom case above shows direct instruction rescuing a failed AI-assisted construction task. | Delayed construction, unfamiliar transfer, learning after ordinary assistant use, and a complete learner-to-reader project remain sparse. | Immediate and delayed unassisted performance, diagnosis, unfamiliar transfer, confidence calibration, and comprehension by the learner's eventual audience. |
| **Accessible and small-screen reading** | Controlled non-AI studies measure low-vision smartphone and blind nonvisual task outcomes. The accessibility sources above expose generated-description failure and organizational scale. | No captured study joins AI-assisted authoring to disabled readers using their own devices and then measures verification, recovery, or harm. | Co-designed alternatives, real devices and assistive technology, task success, errors, time, workload, independent verification, feedback, and remediation. |

### A lifecycle acceptance ledger

The direct answer from the current evidence is deliberately incomplete: no
captured episode follows one AI-assisted visualization from accepted production
through later maintenance, total human-plus-model cost, accessible use on
readers' own devices, a consequential decision, and calibrated trust.

A new adjacent case makes the lifecycle shape more concrete. A [2021–2025
HealthTech visualization program](https://arxiv.org/pdf/2602.23378) followed 21
projects using 84 recorded calls, notes and decision logs, backlogs and issue
trackers, one governance-dashboard case, and a 16-startup survey. The case
reached deployment with patients. Dashboards remained useful when evidence and
responsibility were attached to real decision moments; the authors also warn
that the infrastructure requires continual feedback and maintenance. This is
strong evidence about longitudinal governance visualization. It is not a study
of generative chart authoring, and it does not measure the missing cost,
accessibility, reader-decision, trust, or comparative-maintenance outcomes.

Use one ledger for the evidence, then read the decision layer that belongs to
your role:

| Stage | Evidence to retain | Practitioner | Team or editor | Research and product |
| --- | --- | --- | --- | --- |
| **Frame** | intended audience, decision, source authority, local definitions | Can I explain the claim and reject an incoherent request? | Who owns the metric, editorial judgment, and consequences? | Were the task, stakes, expertise, and baseline declared before use? |
| **Build and correct** | prompts, transformations, direct edits, failures, waits, rollback, rejected output | What did I inspect or repair rather than merely regenerate? | What review changed the artifact, and what remained disputed? | Measure time to first candidate separately from time to accepted work. |
| **Accept and deliver** | named approver, acceptance contract, exact published state, desktop/mobile/access checks | Did the artifact pass the real surface rather than only my preview? | Is approval distinct from publication, and publication distinct from reader use? | Recompute claims and capture the delivered state independently. |
| **Use** | defined reader tasks, comprehension, decisions, confidence, errors, feedback | What did intended readers understand and do? | Did the work support the decision without hiding uncertainty or excluding readers? | Measure decision quality and calibration, not satisfaction alone. |
| **Sustain** | refresh host, dependencies, second maintainer, incidents, correction policy, retirement | Can someone else update, repair, or retire it? | Who owns the next refresh and response when the artifact is wrong? | Follow at least one dependency change, handoff, scheduled update, or incident. |
| **Account** | human time, model or subscription cost, verification, failed generations, delivery, later maintenance | What was the whole cost after the impressive first draft? | Was the gain durable enough to justify adoption? | Compare against current direct work with the same scope and acceptance bar. |

Do not fill a missing cell with evidence from another project. Cases can teach
the acceptance contract collectively; only a row observed from end to end can
establish a complete episode.

## What current products reveal about the direction of travel

Product documentation cannot establish quality, but it shows where vendors are
placing control and context:

- General chat products now combine file upload, executed analysis, interactive
  tables, common chart types, limited direct customization, and download.
- Spreadsheet assistants preserve the installed grid while adding executed
  analysis and charts, but their refresh contracts differ materially: an
  inserted chart may be static or linked to copied rather than original data.
- Artifact builders turn prompts into editable, shareable applications rather
  than returning only prose or a static image.
- Guided editors and visible canvases bind the assistant to explicit settings
  or add new versions beside existing work, making natural language a shortcut
  into a state the user can inspect.
- BI and data-cloud assistants increasingly divide roles: authors maintain
  models, instructions, and permissions; analysts inspect queries and steps;
  consumers ask questions over curated data; administrators monitor usage and
  feedback.
- Open-source systems and skills package declarative schemas, library-specific
  syntax, tests, and production components so a stronger baseline model still
  operates inside a maintained execution contract.
- Research systems pair natural language with direct encodings, visible
  transformed data, history, branches, deterministic perceptual checks, and
  proactive reader scaffolding.

This is a meaningful convergence: natural language is becoming one control
surface inside a structured environment, not the whole environment. As baseline
models improve, generic “make a good chart” instructions are likely to lose
relative value. Maintained business semantics, provenance, deterministic tests,
edit locality, delivery evidence, and reader-specific acceptance are less
likely to be absorbed by the model because they belong to the situation, not
the pretrained baseline.

## A practical way to evaluate a tool or workflow

Do not begin with “Which model made the prettiest chart?” Begin with a real job
and measure the entire path.

1. **Name the environment and audience.** State whether this is exploration,
   governed BI, operational monitoring, public explanation, science, or a
   reusable application. Name who must read or maintain the result.
2. **Choose representative work.** Use self-owned data and include awkward
   semantics, missing values, multi-step transformations, a correction request,
   and the actual delivery surface. Keep a simple task as a control.
3. **Record the baseline.** Measure current human time, error rate, correction
   path, and delivery effort. A tool cannot be said to save time when the
   counterfactual is unknown.
4. **Separate first draft from acceptance.** Record time to first useful
   candidate and time to accepted artifact, including prompt writing, waiting,
   cleanup, verification, and publishing.
5. **Inspect consequential choices.** Can the person see the fields, filters,
   aggregations, transformations, code or semantic query, and generated
   interaction state? Can they make a precise local edit without regeneration?
6. **Test the delivered artifact.** Recompute values, exercise controls, check
   accessibility and responsive behavior, and verify that exported or embedded
   output preserves the intended state.
7. **Test readers separately.** Ask defined readers to state the main claim,
   supporting evidence, uncertainty, and next action. Measure correctness,
   time, confidence, and harmful misreadings. Do not use a model grader as the
   reader sample.

## What remains missing—and what the new evidence partly fills

The gap list is no longer uniformly empty. New controlled studies now quantify
one correction setting, one form of AI-generated chart harm, and several mobile
and nonvisual reader outcomes. A production-workflow interview study adds a
high-stakes dissemination boundary. None closes the central distance between a
controlled task and routine delivered work.

| Question | What can now be said | What remains open | Highest-value next study |
| --- | --- | --- | --- |
| **Can AI critique support work over time on self-owned material?** | Visualizationary followed 13 designers using self-selected data and tools over a three-to-five-day window and found average expert-rated improvement. | observed work was roughly 90–150 minutes per participant, with no baseline, production delivery, or later maintenance | multiweek within-person study on real commissions from intake through publication and update |
| **What does correction cost, and when do people abandon?** | in 108 forced-error task episodes, 7 were not completed and participants prematurely declared completion 31 times; the novice study measured failed repairs; Visualizationary recorded 68 feedback uses and selective rejection | the controlled tasks were engineered to fail and stopped at 15 minutes; there is still no field distribution of correction time, regeneration, rollback, abandonment, or downstream harm | instrument every prompt, edit, wait, undo, verification, and abandonment against a direct-work baseline |
| **Do local semantics improve correctness?** | verification research and current product contracts show why fields, measures, intermediate data, code, and verified queries matter; current BI cases report better trust with narrow curated metrics | no independent head-to-head field test of a general assistant versus the same model with organization-owned semantics | blinded task set with known local definitions, realistic ambiguity, and exact semantic-error scoring |
| **Can the result be delivered and maintained?** | 17 biomedical-visualization practitioners described using AI chiefly for auxiliary tasks and keeping substantial generated imagery out of final scientific work; two 2026 newsroom diaries expose publishing, mobile checks, architecture, proposed updates, and unfinished maintenance; one direct second-maintainer episode exposes a creator-local refresh job; a four-year HealthTech program traces governance visualization through decision work and one patient deployment | the interviews measured judgment rather than delivery; the diaries and handoff are self-reported; the longitudinal program is adjacent governance visualization, not generative authoring, and does not compare maintenance outcomes; no row reaches authenticated generative delivery, regression, and later refresh together | require real publishing, second-person handoff, one dependency change, a scheduled update, and an independently applied acceptance contract |
| **What is the total cost?** | a 108-episode study measured task time and found no significant timing difference among conversational, stepwise, and phasewise interfaces; one 45-minute randomized exercise measured completion; the newsroom diaries add build durations, repeated correction, usage-limit waits, model-cost sensitivity, validation, and remaining work | no common ledger joins human time, model inference or subscription cost, verification, failed generations, later delivery, maintenance, opportunity cost, and abandoned work against a current direct-work baseline | prospective cost diary using time to first candidate *and* time to accepted maintainable artifact |
| **Who benefits, by expertise and work context?** | the human-skills review separates consumption, construction, critique, and connection from data, domain, tool, and delivery resources; the learning chapter adds a positive immediate post-removal comprehension result for proactive scaffolding, adjacent randomized evidence that assisted performance can outrun learning, and ordinary visualization-retention evidence showing faster procedural decay; bounded studies show experienced practitioners filtering critique and using constrained implementation effectively | “novice” remains inconsistently defined; expertise cells are small; delayed construction and far transfer remain mostly unmeasured; no captured longitudinal study causally estimates AI-driven visualization atrophy | preregistered capability-profiled field trial crossing answer-oriented and metacognitive assistance with delayed unassisted transfer in multiple job environments |
| **Does assistance help a reader understand—or harm understanding?** | the 117-person educational experiment found proactive scaffolding outperformed passive Q&A and a data story; a 48-person experiment found lower accuracy with AI-generated misleading charts than with correct charts | both are controlled question-answering tasks; no independent decision quality, public communication, operational action, calibrated trust, or durable field learning | independent reader trials in journalism, BI, education, and public services using delivered artifacts and consequential decisions |
| **Does it work for disabled readers and on small screens?** | a 26-person low-vision smartphone study measured large differences across delivery treatments; a 10-person blind-reader study found modality and chart-type trade-offs across 360 tasks; a 12-person AI-assisted learning study found strong preference and spatial-model benefit for tactile + text + chat, but no measured accuracy lift; a 2025 position paper demonstrates information loss in model-generated description chains | the AI-assisted study was short, small, and did not evaluate delivered artifacts or isolate a chart-vision specialist; broader disability groups, readers’ own devices, screen-reader and keyboard behavior, responsive reflow, verification, and remediation remain open | co-designed AI-versus-non-AI delivery study on readers’ own devices with task success, errors, time, workload, independent verification, recovery, and qualitative experience |
| **What does the work feel like in context—and what are people trying to protect?** | selected first-person accounts now distinguish mandate pressure, first-candidate momentum, repair interruption, reputational and scientific accountability, accessibility, craft, identity, and post-delivery disappointment; underheard-role evidence now includes ordinary spreadsheet allocation, provisional recipients, blind and low-vision learners, and one second maintainer | consequential recipients, disabled readers on their own devices, non-English workplaces, displaced workers, and multiple second maintainers across contexts remain too faint; public accounts cannot estimate prevalence | purposive multi-environment interviews and prospective diaries tied to actual artifacts, usage, corrections, decisions, handoffs, and abandonment |
| **Which scaffolding still helps as models improve?** | proactive questioning beat passive answering in one reader experiment; deterministic perceptual checks and structured interfaces supported bounded authoring; skill projects publish provider-run tests | few independent with-and-without ablations use the same current model, task, and harness, or repeat across model generations | recurring benchmark and field panel that removes one scaffold at a time and records conflicts as well as gains |

Semantic provenance, reader trust calibration, consequential decisions, and
production regression remain especially thin. Harmful misinterpretation now has
one controlled result, but no field estimate. The key experimental unit is not
the generated chart. It is the complete creator-to-artifact-to-reader episode,
with the job, environment, audience, and stakes named in advance.

## Evidence guide: what the cited work actually is

| Evidence | Plain-language description | What it supports here | Important limit |
| --- | --- | --- | --- |
| [2024 State of the Data Visualization Industry](https://www.datavisualizationsociety.org/soti-report-2024) | Online practitioner survey: 980 started, 763 completed, 825 answered the AI-use question; public respondent-level files permit role, task, and incumbent-tool cuts | observed entry points for AI in visualization work and directional differences among this sample | self-selected community survey with item-specific missingness; not representative adoption or product market share |
| [2025 State of Analytics Engineering](https://www.getdbt.com/resources/state-of-analytics-engineering-2025) | dbt survey of 459 practitioners and leaders on development, documentation, and natural-language data use | general versus specialized assistant use and the gap between analytics creation and conversational consumption | vendor-community sample, not visualization-specific, and based on self-report |
| [DWP Microsoft 365 Copilot trial evaluation](https://www.gov.uk/government/publications/an-evaluation-of-dwps-microsoft-copilot-365-trial/an-evaluation-of-dwps-microsoft-365-copilot-trial) (2026 report) | Official mixed-methods workplace evaluation with 1,716 licensed-user survey responses, 2,535 comparison responses, and 19 quota-sampled in-depth interviews | routine work allocation, quiet non-use, Excel-specific limits, conditional trust, habit, time pressure, and stakeholder ownership | nonrandom licence allocation, no pre-trial baseline, self-reported outcomes, and visualization as a subset of a broad office-assistant evaluation |
| [Vibe Visualizing](https://arxiv.org/pdf/2606.08914) (2026 preprint) | Think-aloud study of 20 visualization novices completing 60 ChatGPT chart tasks; 175 charts analyzed; initial prompts replayed on three current multimodal models | novice prompting, design failures, confidence, verification, repair, and changing model behavior | small controlled datasets; cross-model replay was not a live user study |
| [How Good Is ChatGPT in Giving Advice on Your Visualization Design?](https://arxiv.org/pdf/2310.09617) (TOCHI 2025) | Rating study plus 12-practitioner comparison of ChatGPT and human visualization advice | brainstorming value, generic advice, context, expert preference | practitioner sessions used 2023-era GPT-3.5 |
| [How Do Analysts Understand and Verify AI-Assisted Data Analyses?](https://ruoxishang.com/pdf/chi2024_llm_paper.pdf) (CHI 2024) | Design-probe study of 22 professional analysts and 52 verification workflows | how people use code, data, explanations, and visual summaries to verify | prepared tasks at one company; not full chart authoring |
| [Randomized public-health analysis exercise](https://pmc.ncbi.nlm.nih.gov/articles/PMC12521002/) (2025) | 30 analyzed participants assigned to integrated ChatGPT or R/Stata plus ChatGPT for simulated epidemiological work | speed-quality separation, tool integration, serious errors | small, underpowered, simulated 45-minute exercise; no non-AI control |
| [Data Formulator 2](https://arxiv.org/pdf/2408.16119) (CHI 2025) | Mixed-initiative chart-authoring prototype studied with eight corporate participants | direct controls, visible transformations, history, inspection behavior | reproduction tasks, GPT-3.5, no baseline or long-term use |
| [DashChat](https://arxiv.org/pdf/2504.12865) (2025 preprint) | Dashboard-prototyping system with formative interviews and a 28-person evaluation | pre-data prototypes, stakeholder alignment, structured edits | mockups rather than analytical correctness or production delivery |
| [Visualizationary](https://arxiv.org/pdf/2409.13109) (2024 preprint) | Thirteen designers iterated self-selected visualizations with LLM and deterministic perceptual feedback over a three-to-five-day window; three experts rated change | longitudinal critique use, expertise differences, version tracking, and selective rejection of advice | small, short study using GPT-3.5; no critique baseline, production delivery, or later maintenance |
| [GenAI agents and visual-analytics comprehension](https://arxiv.org/pdf/2409.11645) (2024 preprint) | Randomized 117-person comparison of a data story, passive Q&A, and proactive scaffolded dialogue with pre-, intervention-, and post-tests | direct reader-comprehension evidence and the difference between answering and guiding | controlled educational task, limited domain context and comprehension levels, no field deployment |
| [Comparing LLM conversational and graphical interfaces for industrial decision tasks](https://arxiv.org/pdf/2605.31224) (2026 preprint) | Twenty participants used a dashboard and chatbot across four simulated tasks, followed by questionnaires and semi-structured interviews | conversational compression, dashboard overview and auditability, provisional reliance, and the separation between speed and confidence | exploratory proxy sample dominated by computer-science students, short low-stakes tasks, one implementation of each interface, and human-LLM-assisted qualitative coding |
| [AI-supported end-user development for visualization](https://iris.unibs.it/handle/11379/630805) (2025) | Eight interviews in one company followed by three profile-selected, 30-minute design-probe sessions | basic, intermediate, and advanced users' different prompting, transparency, and control needs | exploratory single-company case; three sessions cannot establish usability or efficacy |
| [How Do LLMs See Charts?](https://arxiv.org/pdf/2604.08959) (2026 preprint) | Comparison of 24 people and three multimodal models on 60 synthetic charts | why model and human chart reading are not equivalent | synthetic data and narrow designer-intent labels |
| [Trustworthy by Design](https://arxiv.org/pdf/2503.10892) (2025 preprint) | 37 participants ranking charts and explaining perceived trust | clarity, source, familiarity, integrity, and aesthetics as trust cues | self-reported trust, not comprehension, correctness, or behavior |
| [Vis Ex Machina](https://arxiv.org/pdf/2101.04251) (CHI 2021) | Preregistered experiment with 114 people choosing identical recommendations labeled human or algorithmic | contextual effects of provenance labels | predates generative charts; study charts were researcher-authored |
| [Improving Steering and Verification in AI-Assisted Data Analysis](https://arxiv.org/pdf/2407.02651) (UIST 2024) | Eighteen experienced analysts completed 108 deliberately error-prone tasks across conversational, stepwise, and phasewise interfaces after a 15-person formative study | premature acceptance, non-completion, correction controls, task time, and the cost of structured intervention | forced-error 15-minute tasks using GPT-4 Turbo; no direct-work baseline, publication, or field abandonment |
| [ChartAttack](https://arxiv.org/pdf/2601.12983) (2026 preprint) | Controlled 48-person comparison of correct and AI-generated misleading bar and line charts, with baseline chart-reading phase and adjusted analysis | direct reader-comprehension harm from selected visual misleaders | controlled chart QA and limited chart types; not prevalence, calibrated trust, or consequential decisions |
| [GenAI in biomedical visualization](https://arxiv.org/pdf/2507.14494) (TVCG 2026) | Ninety-minute workflow interviews with 17 biomedical-visualization designers and developers spanning research, production, and dissemination | where high-stakes practitioners use, restrict, or reject AI in real production workflows | purposive Euro-North American qualitative sample; no observed delivery or audience outcomes |
| [Low-vision charts on smartphones](https://ieeevis.b-cdn.net/vis_2024/pdfs/v-full-1917.pdf) (IEEE VIS 2024) | Fourteen formative interviews plus a controlled 26-person comparison of five ways to use bar and line charts on a smartphone | mobile completion, time, error, usability, and workload as delivered-reader outcomes | not an AI-generation study; accurate extraction was assured and participants used an experimenter-provided phone |
| [Sound, Touch, or the Full Monty?](https://pmc.ncbi.nlm.nih.gov/articles/PMC13227614/) (TACCESS 2026) | Ten blind adults completed 360 tasks using screen-reader text, audio-tactile interaction, and a refreshable tactile display across four chart types | modality, device, chart-type, time, accuracy, and workload trade-offs in nonvisual reading | not an AI-assistance study; small controlled sample with limited training |
| [Touching or Chatting](https://arxiv.org/pdf/2607.23065) (2026 preprint) | Counterbalanced study of 12 blind or low-vision participants learning violin plots and clustered heatmaps with tactile + text + GPT-5.2 versus text + GPT-5.2; 263 substantive queries | direct AI-assisted accessibility experience, complementary spatial and conversational roles, preference, and the separation between mental model and accuracy | short small English-speaking US sample; no measured accuracy lift, long-term learning, or delivered-artifact evaluation |
| [Second-maintainer dashboard discussion](https://www.reddit.com/r/dataengineering/comments/1uhidmb/vibe_coded_dashboard_failing_on_a_friday/) (2026 public testimony) | One direct handoff episode plus 91 exposed replies about an AI-built dashboard failing during its creator’s absence | hidden execution host, refresh lineage, semantic documentation, dependencies, ownership, and maintenance expectations | self-selected and unverified testimony; no prevalence, comparison, or causal estimate |
| [Now You See Me](https://arxiv.org/pdf/2602.23378) (2026 preprint) | Embedded 2021–2025 governance-visualization program across 21 early-stage HealthTech projects, 84 recorded calls and operational traces, one dashboard case reaching patient deployment, and a 16-startup survey | decision-linked visualization, evidence reuse, role ownership, iteration, deployment context, and sustainment requirements | embedded UK/EU HealthTech program designed by the authors; not generative visualization authoring and no whole-cost, accessible-reader, calibrated-trust, or comparative-maintenance outcome |
| Public practitioner and reader threads | Unrecruited discussions about current BI and spreadsheet work and chart reactions, including a 190-entry Excel discussion and a July 2026 dashboard-versus-chat discussion | concrete jobs, delights, friction, abandonment, changes in peer learning, trust repairs, and reader questions that occur in practice | self-selected testimony; establishes possibility of an experience, never prevalence or a population trend |
| [Selected first-person creator logs](https://sef.kloninger.com/posts/claude-dataviz/) and [practitioner interviews](https://www.datacamp.com/podcast/how-next-gen-data-analytics-powers-your-ai-strategy) | Named accounts spanning early experimentation, guided learning, personal dashboards, visualization coursework, analytics practice, data journalism, accessibility, and analog craft | how anticipation, delight, frustration, responsibility, refusal, and professional identity change across stages of work | editorially selected and unusually articulate accounts; self-reported episodes are not comparative tests, population estimates, or independent artifact audits |
| [Two 2026 Reuters Institute build diaries](https://reutersinstitute.politics.ox.ac.uk/news/weeks-work-days-how-i-rebuilt-two-data-journalism-projects-ai) | First-person accounts of rebuilding two data-journalism projects and constructing one national health dashboard with current coding agents | production sequence, time compression, visible defects, correction, architecture, mobile checks, validation, ownership, and unfinished work | three projects built and checked by their authors; no equal-budget baseline, independent audit, reader outcome, or maintenance follow-through |
| [Playing Telephone with Generative Models](https://arxiv.org/pdf/2508.12192) (2025 position paper) | One source visualization passed through four GPT-4o and Claude Sonnet 4 description-and-image chains | mechanisms of information loss, fabrication, misplaced emphasis, poor design comprehension, under-description, overconfidence, verification disability, and compelled reliance | authors explicitly frame the work as a provocation rather than a controlled experiment; no disabled-reader study or error-rate estimate |
| [Doom or Deliciousness](https://www.mcnutt.in/assets/doom-n-fruit.pdf) (CGF 2023) | Semi-structured interviews with 21 experts in visualization or HCI, art or art history, and machine learning | early field expectations about help and harm across datafication, transformation, visualization, and interaction | convenience sample, mostly limited direct generative-tool experience, image-generation emphasis, and speculative framing |
| Dated field essays, interviews, and teaching or design cases | Role-diverse assessments from visualization researchers and educators, analytics and library practitioners, a blind journalist, a creative studio, and data journalists | purposes, workflow mechanisms, productive disagreements, lived constraints, and questions the controlled literature should test | perspective and self-report are not prevalence, comparative efficacy, or reader-outcome evidence; positionality and publication date remain attached |
| Current product and project documentation | Provider or maintainer descriptions for general assistants, spreadsheets, notebooks, editors, governed BI, data clouds, open-source systems, and skills | the feature, data-access, target-user, control, and delivery contracts available in August 2026 | self-reported feature scope and positioning, not independent adoption or comparative performance |

## Scope and limits

This review is a moment-in-time synthesis, not a market-share study or product
ranking. Twenty-two human studies and structured workplace evaluations were read
in full, along with practitioner surveys, current provider and open-source
documentation, public practitioner and reader discussions, first-person
production cases, and dated essays and interviews from several professional positions. Study
populations, tasks, models, and evidence vintages remain visible because results
are not directly interchangeable.

Public threads were used to discover and illustrate experiences that controlled
research often omits. They were not sampled to estimate prevalence, and quoted
claims were not independently reproduced. Vendor documents establish what a
product says it supports, not that the feature is accurate or useful. Several
important human-computer-interaction studies use 2021–2024 systems; their raw
model comparisons are dated, while their observations about context,
verification, and human control remain relevant hypotheses for current tools.

The strongest counter-reading is that better 2026 models may erase many errors
reported in earlier work. The fresh novice study partly supports that: replayed
prompts produced fewer design flaws with newer systems. It also found new
failure modes from richer outputs, including render failures, misleading
interpretations, latency, unclear interaction state, and greater verification
burden. Capability is moving; the location of the human work is moving with it.

## Update log

- **2026-08-15 — Lifecycle evidence and audience decision layer.** Added a
  four-year HealthTech governance-visualization program, preserved why it is
  adjacent rather than generative-authoring evidence, and introduced a shared
  lifecycle acceptance ledger translated for practitioners, leaders, editors,
  researchers, product teams, and accessibility work.
- **2026-08-14 — Initial public edition.** Mapped practitioner jobs and tool
  choices, synthesized creator and reader accounts, and separated possible
  output, accomplished work, delivery, and lived outcome.
