The practitioner and reader experience of AI-assisted data visualization
Status: research snapshot, evidence cut 2026-08-16. Recheck by 2026-11-16, or earlier after a material change in the major general-purpose, visualization, or business-intelligence assistants discussed below.
Companions: The state of AI-assisted data visualization research, Human skills and banked gains in AI-assisted data visualization, and Learning data visualization when AI can make the chart.
This report asks a different question from the research review. Instead of starting with architectures and benchmarks, it asks what it is like to use AI to make or read a data visualization in August 2026. It covers the jobs people are trying to perform, the kinds of tools available, measured studies of human behavior, and public accounts of what feels remarkable or frustrating.
The short answer is that AI is already good at getting someone from a blank page to a plausible first thing. It is much less reliable at carrying the work through local meaning, correction, delivery, and audience understanding. That distinction explains why a demo can be astonishing, a practitioner can save real time, and the finished chart can still be wrong or unhelpful.
Executive summary
AI-assisted visualization has three different success thresholds:
- Possible: a model or tool can produce a chart, dashboard, analysis, or interactive artifact under some conditions.
- Accomplished: a practitioner can turn that output into correct, inspectable, editable, and deliverable work at an acceptable total cost.
- Experienced: creators understand and can control the process, while readers understand the result, notice its limits, and place appropriate trust in it.
Most public demonstrations stop at the first threshold. The strongest research systems reach into the second by exposing transformed data, code, semantic models, edit history, or direct controls. Very little work reaches the third: independent reader comprehension, decision quality, calibrated trust, accessibility, and mobile use remain sparsely measured.
Across the evidence, five conclusions are reasonably stable:
- The reliable win is compression, not delegation. AI can accelerate first drafts, chart reproduction, data preparation, routine code, documentation, and prototype variations. Those gains are real even when a human retains the analytical claim and final design.
- Speed and polish are unsafe quality signals. In a randomized data-analysis exercise, the faster integrated-AI group did not score better and produced fewer serious-error-free submissions. In a 2026 novice study, confidence and satisfaction remained fairly high despite pervasive chart flaws, task noncompliance, and some incorrect interpretations.
- Context and inspectability determine whether a draft becomes usable work. Systems do better when they can see governed metric definitions, field meanings, intermediate tables, code, worksheet state, or a restricted visual specification. Practitioners verify through artifacts they already understand; a mysterious result is difficult to trust or repair.
- Expertise changes the experience rather than simply increasing the gain. Experts can reject nonsense, spot broken assumptions, and redirect a model. They also recognize when writing the prompt, waiting for a regeneration, or cleaning the output costs more than direct work. Novices gain access but may be least equipped to notice consequential errors.
- A good creator experience and a good reader experience are different outcomes. Fast generation, visual appeal, and creator satisfaction do not establish that readers understand the intended claim. Readers use clarity, familiarity, sourcing, and aesthetics as trust cues, but those cues do not establish correctness.
A sixth conclusion is becoming visible but is less settled: adoption is entering through existing work rather than replacing the visualization stack. In the best available practitioner survey, AI use clustered around preparation, analysis, coding, and ideation, while Excel remained the most widely used visualization tool. Current BI testimony similarly describes chat taking ad hoc questions while recurring dashboards retain shared definitions. This is a directional reading of self-selected evidence, not a market-share estimate.
The practical implication is not “use AI” or “avoid AI.” Use it where the work is reversible, the output remains inspectable, and the human can recognize a bad result. Demand stronger grounding, deterministic checks, and explicit reader testing as the semantic stakes, delivery costs, or audience consequences rise.
The gap that organizes the evidence
The same output looks different at each threshold. A generated dashboard can render and still fail as work; a correct chart can still fail its reader.
| Threshold | The question being answered | Evidence that counts | Common false substitute |
|---|---|---|---|
| Possible | Can the system produce the requested class of output at least sometimes? | executable code, a rendered chart, task completion, supported feature contract | a polished screenshot or vendor description |
| Accomplished | Can someone finish the actual job correctly and maintain the result? | correct data and transformations, inspectable decisions, correction cost, delivery, reuse, total time and cost | first-pass speed, nonblank rendering, or a benchmark score alone |
| Experienced — creator | Can the person understand, steer, correct, and appropriately trust the process? | observed behavior, successful repairs, calibrated confidence, workload, abandonment, learning | satisfaction or stated preference alone |
| Experienced — reader | Does the final audience understand the claim and its limits and make an appropriate judgment? | comprehension, recall, decision quality, trust calibration, accessibility, behavior on the delivered device | model chart-reading, aesthetic ratings, or author confidence |
Three controlled bridges connect capability to human consequence
The gap is not empty. A bounded audit of ten primary cases finds three partial bridges that keep technical and human evidence on one controlled path:
| Study | What remains joined | Where the bridge stops |
|---|---|---|
| HAIChart | automated recommendation results and 17-person comparative use over the same eight KaggleBench datasets | no immutable configuration, frozen acceptance contract, or field delivery; the system is a reinforcement-learning recommender, not a current LLM authoring assistant |
| Interactive task decomposition | system errors, interventions, and correction behavior within 108 experienced-analyst episodes | no direct-work baseline, ordinary field distribution, accepted delivery, or later use |
| ChartAttack | generated misleading-chart instances carried into a 48-person reader experiment | no prevalence estimate, consequential decision, real delivery context, later correction, or whole cost |
That is three controlled partial bridges, zero exact-version accepted-delivery bridges, and zero complete episodes in the ten audited cases and named search surfaces through 15 August 2026. It rejects both an empty-bridge story and a synthetic end-to-end story assembled across studies.
This is not a maturity ladder on which every visualization must climb. A throwaway sketch may only need to be possible. A board metric, public-health chart, or published explanatory graphic needs a defensible path through all four rows.
Who is actually reaching for what
The evidence is better at showing where AI has entered the workflow than at naming a winning product. The 2024 Data Visualization Society survey was self-selected but unusually useful: 980 practitioners started it, 763 completed it, and 825 answered the AI-use question. Of those 825, 305 (37.0%) said they had used AI in visualization work during the prior year, up 13 percentage points from 2023; 492 (59.6%) said no and 28 (3.4%) were unsure.
Among the 305 AI users, multiple selections were allowed. The work was weighted toward the beginning of the lifecycle:
| Job reported by AI users | Respondents | Share of 305 AI users |
|---|---|---|
| Prepare or clean data | 184 | 60.3% |
| Analyze data | 134 | 43.9% |
| Ideate or storyboard | 119 | 39.0% |
| Produce visualizations | 92 | 30.2% |
| Support a visualization team | 27 | 8.9% |
This is not a story of visualization specialists abandoning incumbent tools. Excel was used often or sometimes by almost identical shares of AI users and nonusers in the public microdata (66.9% versus 66.7%), and Tableau use was also similar (43.9% versus 47.4%). AI users were more likely to report Python (42.6% versus 27.4%), D3 (30.5% versus 13.4%), Observable (16.1% versus 5.7%), Figma (40.0% versus 28.5%), and Datawrapper (19.7% versus 10.2%). The defensible interpretation is that early adopters in this sample often already crossed code, design, and publishing environments. Co-use does not show which tool caused adoption.
A second, non-visualization-specific survey sharpens the split. In the 2025 State of Analytics Engineering, 80% of 459 respondents reported day-to-day AI use, up from 30% one year earlier. Seventy percent used AI for analytics development, mainly through general assistants such as ChatGPT, Claude, and Gemini; roughly 25% used specialized AI inside development tooling. Only 30% were using natural language to consume data, while another 29% wanted to and 23% had experimented. The sample came from a vendor community and is not visualization-specific, but it reinforces a plausible sequence: code and documentation adoption precede trusted conversational consumption.
The audience-to-environment map
The table separates four kinds of evidence that are often blurred together: surveyed behavior, public cases, provider-defined target users, and our inferred fit from the work surface.
| Who | What they appear to reach for first | Job they are trying to do | Status of the evidence |
|---|---|---|---|
| Visualization practitioners with an existing stack | a general assistant beside Excel, Tableau, Python, R, D3, Figma, or a publishing tool | clean data, debug code, learn a technique, storyboard, draft labels, make a first view | Observed survey pattern. The DVS sample measures AI jobs and incumbent-tool co-use, not specific AI-product share. |
| Analytics engineers and code-capable analysts | ChatGPT, Claude, Gemini, Copilot-style code help, then project-aware notebook or development agents | write and explain SQL or Python, document models, debug pipelines, generate a chart inside an analysis | Observed survey pattern. The dbt sample is vendor-adjacent and broader than visualization. |
| Spreadsheet-native analysts, operations teams, and finance users | AI inside Excel or Sheets; newer AI spreadsheets when code or live connections outgrow a grid | ask about a range, create formulas and charts, preserve familiar cells, hand work to colleagues | Installed-base inference plus provider target. Excel is widely observed; adoption of its AI features was not measured here. |
| BI authors and data teams | AI inside Power BI, Tableau, Looker, Hex, Sigma, or a governed data platform | author reports, define or reuse measures, inspect queries, answer stakeholder follow-ups, govern access | Provider target plus public cases. Independent head-to-head field evidence remains absent. |
| Business consumers and executives | conversational layers such as Genie, Spotter, Ask Sigma, Qlik Answers, Amazon Q, or Oracle’s assistant | retrieve a metric, ask why it changed, get an ad hoc cut, avoid navigating a complex report | Provider target with self-selected testimony. Current cases say chat complements recurring dashboards and loses trust quickly without curated definitions. |
| Designers, journalists, and explanatory communicators | general assistants for ideation or code, then Figma, Datawrapper, Flourish, or a newsroom-owned production stack | explore, create variants, annotate, improve accessibility, implement a story without surrendering editorial judgment | Survey subgroup and workflow inference. Journalism had 24 AI users among 47 valid DVS responses, too small and self-selected for a population rate. |
| Developers and data-app builders | coding agents, generated applications, declarative chart specifications, open-source agent frameworks, and reusable skills | make bespoke interaction, integrate data and UI, automate repeatable production, deploy an application | Project positioning and inferred fit. Repository activity shows availability, not routine user success. |
The July 2026 public BI discussion helps explain the business-consumer row without estimating its size. Reported deployments use chat for exploration, follow-up questions, and requests that previously became analyst emails; recurring dashboards still provide shared monitoring and durable definitions. Participants repeatedly tied trust to a narrow mart, approved metrics, visible refusals, and an owner for each measure. Several described slow or failed pilots, and one reported that a single wrong executive number ended trust. These are contemporary cases, not prevalence evidence.
The jobs: visualization is a workflow, not a prompt
“Make a chart” hides work before, during, and after visual encoding. Separating the jobs matters because the same tool can be strong at one and harmful at the next.
| Phase | Job | Useful assistance | Human responsibility that does not disappear |
|---|---|---|---|
| Define | Frame the question | restate a request, propose hypotheses, identify missing information | decide what matters, tie the work to a decision, reject an incoherent request |
| Define | Acquire and govern | find likely tables, draft access requests, describe schemas | choose authoritative sources, respect permissions, establish ownership and allowed use |
| Work the data | Prepare and transform | clean types, join, reshape, write SQL or Python, create routine calculations | define measures, resolve ambiguous fields, inspect exclusions and missingness |
| Work the data | Analyze | calculate summaries, fit standard models, compare groups, expose anomalies | choose a valid method, interpret uncertainty, know which causal claim is unavailable |
| Work the data | Explore | generate many cheap views, suggest cuts, translate follow-ups into operations | recognize spurious patterns, investigate causes, know when to stop branching |
| Make the artifact | Sketch or reproduce | remove blank-page work, copy a reference, create dashboard mockups and variants | decide whether the reference belongs in this context and what must change |
| Make the artifact | Choose and encode | suggest chart forms, mappings, aggregations, annotations, and emphasis | make the analytical and rhetorical choice for this data, claim, and audience |
| Make the artifact | Critique and refine | flag common perceptual problems, apply repetitive settings, draft labels and accessibility text | resolve composition, tone, brand, direct-manipulation details, and conflicting advice |
| Make the artifact | Verify and debug | expose code and intermediate tables, generate tests, replay interactions, compare results | state expected behavior, choose decisive checks, judge local semantic correctness |
| Deliver | Package and publish | assemble a report or app, write documentation, export standard formats | accept the delivered artifact, satisfy privacy and review gates, test the actual surface |
| Deliver | Explain and interrogate | provide tooltips, summaries, guided questions, and grounded conversational follow-ups | preserve the source claim, communicate uncertainty, evaluate reader understanding |
| Sustain | Maintain and update | regenerate routine artifacts, flag drift, propose regression checks | own dependency changes, semantic revisions, handoff, correction, and retirement |
The public cases are most consistent about work adjacent to the final visual: SQL help, calculations, documentation, mockups, tooltips, bulk changes, and alternative ideas. Claims that a system “built the dashboard” often compress substantial context: an existing screenshot, prepared source tables, a semantic model, a known template, or extensive cleanup.
Listening to the field: what this work actually feels like
The accounts below are selected, not sampled. They come from first-person project logs, interviews, public workplace discussions, and one qualitative study of biomedical-visualization practitioners. Public handles, roles, and outcomes are self-described unless a named publication or study says otherwise. The accounts demonstrate that an experience was reported; they do not estimate how common it is or prove that the artifact worked as described.
One distinction makes the listening more useful: attitude and behavior are not the same dimension. A person can be delighted and unwilling to ship, anxious and still using the tool, skeptical but grateful for boilerplate help, or enthusiastic while retaining direct inspection and repair. A 17-person biomedical-visualization interview study observed five such modes, from enthusiastic adoption through skeptical avoidance. That specialized typology is not imposed on every account here. It is a reminder not to reduce the field to boosters and opponents.
Before the work: invitation, pressure, and hesitation
- An analytics team lead described leadership demanding AI “thought leaders” even though an end-user chatbot was unreliable and a pivot table was faster. The frustration was not simply with the model. It came from having to perform AI adoption before identifying a useful job.
- Visualization coder Jisell Howe initially feared that instant code would remove the rewarding journey from idea to customized chart. A concrete project changed that view: wrong syntax remained, but faster search and troubleshooting restored creative momentum.
- A master’s student with little Python experience began a guided practical unsettled by programming. AI made a bar chart quickly; an instructor then supplied the design knowledge needed to improve it. The felt gain was access and accomplishment, not independent mastery.
- Biomedical-visualization specialists included eager adopters, interested non-users, cautious limiters, and principled avoiders. Resistance sometimes reflected scientific risk, threatened livelihoods, or pride in cultivated craft rather than unfamiliarity with the technology.
The first candidate: access, velocity, and the rush of possibility
- A Japanese independent creator spent two days revising a personal dashboard and joked that the model revised their professional dignity too. A loose four-line instruction worked because the current files, desired fixes, and URLs were available. The surprise was that organized material mattered more than a polished prompt.
- A technologist taking a visualization class described a week with an agent as “plain fun”. It helped process a 500 GB dataset, build tests, produce an exploratory dashboard, debug operations, and draft a presentation. The delight came from uninterrupted momentum across a whole project, not from one chart appearing.
- One experienced BI analyst said AI compressed an unfamiliar API task from an estimated week to a day, then cautioned that prior proficiency probably made that leverage possible. Another novice reported a much smaller but still meaningful gain: writing cleaning and validation scripts they could not otherwise have produced.
The working loop: babysitting, interruption, and selective handback
- In one Power BI discussion, generated greenfield structure arrived quickly, while stale filters, bookmark identifiers, whitespace churn, and slow edit-preview cycles made local refinement painful. A practitioner called the approach good for new reports and clumsy for small changes; an experienced engineer said describing the change could cost more than doing it.
- Analytics practitioner and educator Christina Stathopoulos described a useful analytical “sidekick” that still needed babysitting. In one episode it generated a chart and then stated the opposite of what the chart showed. The feared outcome was not an abstract benchmark failure but looking careless in front of stakeholders.
- Journalist Jaemark Tordecilla saw a recognizable site appear in about an hour. Local geography, evidence notes, usage limits, mobile behavior, architecture, and fact-checking then consumed repeated rounds. The work became most demanding exactly where the first candidate looked closest to done.
Acceptance: reputation, accountability, access, and the right to refuse
- A dashboard consultant watched an agent reconstruct roughly a week’s work from the finished screenshot and source tables in about 25 minutes. The result prompted a blunt professional question: what is the consultant now charging for—the wrong portion, the ability to find it, or the invisible work already embedded in the reference?
- Biomedical-visualization practitioners welcomed boilerplate, inspiration, translation, and nontechnical assets but drew firm boundaries around anatomy, patient education, final images, and explainability. Some did not want to automate apparently mindless rendering because that was also where flow, control, and creative joy lived.
- Blind journalist Johny Cassidy described exclusion from charts with missing or useless descriptions, then the relief of colleagues making access a shared habit. The experience makes clear why automatically emitted alt text is not yet accessible delivery: readers also need useful alternatives, feedback, and accountable correction.
- When public readers encountered a dramatic chart about AI-generated content, some accepted its story while others challenged its classifier, denominator, symmetry, and implied substitution. The discussion became an argument about AI itself before the chart’s method was resolved. Readers see a claim, not the creator’s prompts and repairs.
After delivery: use, learning, identity, and whether the saved effort mattered
- In the consultant discussion, another BI practitioner recalled an urgent dashboard that a manager praised and distributed to six colleagues. Three months later the usage record showed one manager view. Faster production can create an unused artifact faster; creator pride, stakeholder praise, and reader use are different outcomes.
- A BI developer described using AI not to generate the final chart but to document SQL and calculations, review junior work, and rehearse stakeholder questions. The durable value was work around the dashboard that made it easier to explain and maintain.
- Freelance data journalist Emilia Ruzicka deliberately collected and drew personal data by hand. The slow, imperfect practice produced attention, experimentation, and a different relationship with the data. It is a reminder that not every friction is waste.
- In the leadership-pressure discussion, one commenter said their data team first feared automating itself away and later became the maintainer of the definitions and systems everyone relied on. Another described a leader trusting AI blindly and pushing bad recommendations back onto the team. Automation can elevate context work—or transfer more defensive work to the people who already own it.
Listening past the original creator
A second listening pass pursued people who were mostly absent above: routine spreadsheet workers, quiet non-users, recipients of generated analysis, and the people expected to live with the result. It changes the picture in ways a collection of creator success stories cannot.
| Setting and voice | Reported episode | What it makes visible | Limit |
|---|---|---|---|
| Public-sector staff using Microsoft 365 in ordinary work | A DWP mixed-methods evaluation found that people routed work according to expertise, time, trust, data sensitivity, and habit. Experienced or data-heavy work often stayed in established methods; Excel users specifically reported trouble with complex data, charts, and formatting. One person valued the assistant yet was too busy to remember to use it consistently. | Non-use can be local and quiet. A person may value the overall assistant while declining it for the spreadsheet, chart, or stakeholder-owned task in front of them. | 1,716 user survey responses, 2,535 comparison responses, and 19 interviews provide substantial workplace evidence, but licences were not randomized, no pre-trial baseline existed, outcomes were self-reported, and visualization is a subset. |
| A supply-chain analyst and a spreadsheet-help community | A self-described intermediate Excel user said company encouragement made AI the first stop for reports and reduced visits to a peer forum. The 190-entry discussion ranged from useful Power Query help to invented functions, wrong references, damaged formulas, and decisions never to ask again. Several people worried about losing the learning that comes from solving another person’s problem in public. | Assistance changes the support ecology, not only task time. Instant private answers can lower access cost while thinning public explanation, correction, and incidental learning. | One self-selected thread establishes that these experiences and concerns were reported. It cannot establish prevalence or a forum decline caused by AI. |
| People making simulated industrial decisions | Twenty participants used both a dashboard and an LLM conversational interface. They experienced conversation as compressed retrieval and mental offloading, but often regarded the result as provisionally useful rather than defensible. When speed was the criterion, 15 chose the chatbot and 3 the dashboard; when confidence was the criterion, all 18 decisive responses chose the dashboard. Several wanted both: overview to discover the question, conversation to retrieve or assemble the answer. | Usefulness and reliance separate for recipients just as delight and acceptance separate for creators. Extra visual interaction can be valuable when it exposes the values and relationships needed to defend a decision. | The exploratory study used four short simulated tasks; 18 of 20 participants had computer-science backgrounds and were proxies rather than experienced industrial decision makers. |
| Program recipients and an observer of an executive sales dashboard | A social-impact dashboard case reports that program managers, evaluators, and funders wanted overview and drill-down, definitions, neutral language, comparative context, and review support. In a separate named public critique, telling an executive to verify regenerated charts by checking raw data felt like transferring the product’s verification burden to the person seeking a quick decision. | A recipient does not simply need a shorter summary. The path from alert to explanation and from claim to evidence must be proportionate to the recipient’s actual job. | The design case omits its interview count and instruments; the critique is one observer’s account and does not establish that the output was wrong or how the intended executive reacted. |
| Blind and low-vision people learning unfamiliar chart forms | In a 12-participant study, people learned violin plots and clustered heatmaps with either tactile charts plus text and GPT-5.2 or text and GPT-5.2. Eleven preferred the tactile combination; one said the better mode depended on chart complexity; none preferred text and chat alone. Participants described touch as supplying the spatial model and chat as supplying flexible clarification. | Accessibility is not one generated description. Complementary modalities can answer different needs, and an open-ended assistant cannot always detect misunderstanding or supply questions a novice does not know to ask. | Preference and reported spatial understanding improved, but measured chart-understanding accuracy did not. This was a short study with a small English-speaking US sample, not a long-term learning or delivered-work evaluation. |
| A second maintainer inheriting an AI-built dashboard | A public discussion begins with a dashboard failing while its creator was away. The inheriting maintainer found that refresh and scheduling depended on the creator’s laptop and replaced the setup with a proper pipeline. Replies described missing source and metric semantics, hidden dependencies, absent comments and documentation, and an implicit service expectation nobody had owned. | The maintenance boundary begins before handoff. Execution host, refresh lineage, definitions, dependencies, documentation, ownership, and rollback must be inspectable while the creator is still present. | This is one self-selected, unverified public episode plus discussion. It establishes that the experience was reported, not how common it is or whether AI rather than ordinary software practice caused every defect. |
Quiet withdrawal is therefore a gradient: task-level refusal, habitual fallback, feature abandonment, substitution of one model for another, and provisional use without confident reliance are different behaviors. A retained licence or active session can hide all five.
The direct second-maintainer gap is no longer empty, but it remains extremely thin. One episode is enough to make the hidden execution host and ownership boundary concrete; it is not enough to estimate frequency, compare AI-assisted and direct work, or describe what happens across organizations and tool stacks.
What these voices reveal that aggregate capability evidence does not
- Delight is often momentum: the pleasure comes from staying in motion across chores, tools, and formerly blocked work—not necessarily from a better final chart.
- Frustration is often interruption: assistance becomes costly when a fluent practitioner must stop direct work to specify, wait, inspect, and repair a small change.
- Fear has distinct objects: job loss, reputation, lost craft, weakened learning, inaccessible delivery, scientific harm, and vendor dependence are not interchangeable concerns.
- Trust is personal: practitioners picture the stakeholder who will spot an obvious mistake, the patient affected by a wrong image, the client whose work bears their name, or the reader who cannot independently test a description.
- Refusal can be expertise: declining to use AI for a final visual may be a reasoned response to consequence, accountability, or the value of doing the supposedly low-level work directly.
- The afterlife remains faint, not empty: one direct handoff failure now exposes hidden execution and ownership assumptions, but routine reader use, maintenance across releases, correction, and retirement remain largely unobserved.
What people in the field say has changed
Benchmarks tell us whether a system can perform a defined task. The perspectives below explain why that task matters, what surrounds it in practice, and where people disagree about the purpose of assistance. They include first-person build diaries, practitioner essays, interviews, design-studio reflections, a classroom exercise, and two research papers. They establish that an experience or argument exists. Except where a measured study is named, they do not establish frequency, comparative performance, or a general causal effect.
Across these selected perspectives, implementation is getting cheaper. The substantive disagreement is what the saved effort should buy.
| Tension | What one perspective contributes | What the counterweight contributes | What context should decide |
|---|---|---|---|
| Construction speed versus verification debt | Two 2026 newsroom diaries describe work that once took weeks appearing in days or hours. They also report incorrect totals, misplaced geographies, missed notes, repeated correction, usage limits, mobile checks, and new validation machinery. | Analytics practitioner Benn Stancil argues that a plausible chart cannot validate the calculation behind it; local definitions, executable lineage, and review remain necessary. | Semantic stakes, independence of the check, cost of a wrong result, and whether the complete pipeline—not only the first render—is owned. |
| Lower implementation barriers versus more valuable judgment | Enrico Bertini maps opportunities across acquisition, wrangling, analysis, data tours, and creative exploration. Alberto Cairo treats easier code as capacity extension. | Both make question quality, synthesis, interpretation, and final analytical choice more important rather than less. Cairo also distinguishes exploratory, explanatory, essayistic, artistic, and poetic visualization purposes. | Purpose, audience, consequence, the person’s capability profile, and whether the resulting claim can be inspected and challenged. |
| Frictionless production versus productive friction | Removing syntax and formatting work can let a person reach a candidate before the question goes cold. | In one classroom exercise, students spent 25–30 minutes failing with AI before direct instruction supplied the missing chart structure. An analog data-journalism reflection describes slow, manual work as a source of attention, experimentation, and care. | Whether the effort is obsolete mechanics or practice in structure, skepticism, memory, and meaning; whether assistance will later be withdrawn. |
| Automated access versus independent verification | Generative systems can draft descriptions and alternative representations at a scale that newsrooms have struggled to support. | Elavsky and Xiong Bearfield show how model-to-model chart descriptions can lose data, source, uncertainty, design choices, and authorial purpose. Blind journalist Johny Cassidy describes access as an organizational, multimodal, and reader-feedback problem—not an alt-text checkbox. | Whether the reader can test the account, reach the underlying data or authorial description, choose another modality, report a failure, and obtain remediation. |
| Static charts versus conversational or divergent representations | Richard Brath asks which visualizations remain useful when a model can extract and organize insights. Elijah Meeks proposes making audience, intent, suggestion, interpretation, and conversation first-class framework objects. | Domestic Data Streamers argues that visualization remains a distinct language for comparison, human experience, and creative divergence that a textual answer does not replace. | Reader task, need for overview or a shared evidence surface, qualitative meaning, desired novelty, and whether interaction leaves an auditable trace. |
An early formal baseline helps distinguish current observation from hindsight. In Doom or Deliciousness, 21 visualization, HCI, art, and machine-learning experts interviewed in 2023 anticipated rapid prototyping, personalization, and creative assistance alongside bias, unreliable output, inattention, lost agency, and unearned trust. The study used a convenience sample, foregrounded image-generation systems, and asked participants to think beyond the technology they had used. It is useful as a “what aged well?” comparison, not a verdict on current systems.
Two current build diaries show where the time moved
The two journalism cases are unusually detailed accounts of current production. They are still self-reports by the people who built and checked their own work. Keeping the gain beside the incurred work prevents “built in a week” from becoming a quality claim.
| Lifecycle stage | Ottaviani: two reconstructed projects | Tordecilla: national health dashboard | What neither case establishes |
|---|---|---|---|
| Baseline and first candidate | One map was rebuilt in two days after taking roughly three weeks in 2012; a second portal was attempted in a much more autonomous pass. | A repository became a recognizable site in about an hour; the full dashboard, maps, charts, analysis, and checking workflow took one week. | An equal-budget comparison with the best current direct workflow or another practitioner. |
| Defects and correction | Incorrect cluster and chart totals and incomplete bilingual text required source comparison and several correction rounds. | Highly urbanized cities rendered in the sea or in broken shapes; source notes and exact evidence were sometimes missed; long runs took shortcuts. | A field distribution of error, correction time, regression, premature acceptance, or abandonment. |
| Architecture and control | The author decomposed the work into about 20 tasks, ordered foundational work first, and kept code in an open repository. | The author separated extraction, data, presentation, insights, and fact-checking so a source error could be repaired and replayed. | That decomposition or a model-built validator independently improves correctness across projects. |
| Delivery and readers | The projects were published as working prototypes; the author identifies attention, understanding, and civic use as the remaining bottleneck. | Mobile rendering was checked and stakeholders reacted to the tool, but no reader-comprehension or decision study was conducted. | Audience understanding, accessibility, consequential use, or sustained adoption. |
| Sustainment | Proposed next steps include record-by-record reconciliation and scheduled updates. | Refactoring and longer-term maintenance remained after the one-week build. | Later refreshes, dependency changes, second-person handoff, incident response, or retirement. |
The tools: a showcase of the current approaches
There is no single AI-visualization product category. A useful comparison asks five literal questions: What can the assistant see? What can it change? Which intermediate state can the user inspect? How local is a correction? What form is actually delivered? The examples below are a current landscape, not a ranking. Product pages establish feature contracts and intended users, not independent quality or adoption.
General assistants and specialist file chat
| Tool | Who it is positioned for | What it makes or changes | Where control returns to the person |
|---|---|---|---|
| ChatGPT data analysis | anyone bringing a bounded file or table into a conversation | cleaning, executed analysis, common static or interactive charts, explanations, downloadable results | inspect generated code and tables; supply local definitions and move production work elsewhere |
| Claude Artifacts | people who want a shareable custom visual or small application | generated HTML or application code, interaction, explanatory interfaces | inspect and edit the artifact; assume responsibility for state, accessibility, hosting, and maintenance |
| Julius | file- or connection-oriented analysts seeking a dedicated data chat | analysis, charts, statistical models, reports, and presentations | continue through prompts or exported artifacts; independent repair and correctness evidence is sparse |
Spreadsheet-native assistance
| Tool | Who it is positioned for | What it makes or changes | Important contract boundary |
|---|---|---|---|
| Copilot in Excel | Excel beginners through advanced analysts | Python-backed answers and optional static chart or table insertion; an advanced mode can create refreshable Python cells | the direct-answer mode does not modify the workbook; inserted charts are static and non-refreshable |
| Gemini in Google Sheets | people already working in a shared spreadsheet | summaries, formulas, prompted edits, and editable charts inserted with supporting data on a new tab | the generated chart links to the new tab, not to subsequent changes in the original dataset |
| Bricks | teams wanting spreadsheets, dashboards, slides, and collaboration in one workspace | grid-based analysis, live dashboard boards, presentation material | combines surfaces to reduce handoff, but current evidence is provider documentation rather than comparative use |
| Quadratic | spreadsheet users who also need Python, SQL, JavaScript, or live database connections | AI-generated code and queries, direct grid edits, code cells returning results to the sheet | exposes code and schema inside a familiar grid; asks users to manage a more technical workbook |
Notebooks and visible canvases
| Tool | Who it is positioned for | What it makes or changes | Important contract boundary |
|---|---|---|---|
| Hex AI | technical notebook authors, data teams, and consumers of published data apps | SQL, Python, chart, pivot, and Markdown cells; generated apps; semantic models; consumer chat | roles are explicit: technical users audit code, consumers query curated apps, managers govern semantic context |
| Observable Canvases AI | collaborative analysts who want AI work beside their own | new tables, transformations, SQL, and charts placed on the current canvas | the assistant adds new versions rather than silently editing or deleting; its view is bounded by viewport and schema |
Guided storytelling editor
| Tool | Who it is positioned for | What it makes or changes | Important contract boundary |
|---|---|---|---|
| Flourish AI | beginners, visualization practitioners, and teams refining a chart in a known editor | native editor settings for styling, labels, sources, annotation, and accessibility | proposed changes remain reversible in the editor; the assistant does not edit source data or invent missing intent |
Governed BI and data-cloud assistants
| Tool | Primary surface | Intended reach | Distinguishing control or dependency |
|---|---|---|---|
| Power BI Copilot | report authoring inside Power BI | BI authors creating or modifying report pages | grounded in the model and report state; quality depends on maintained measures, descriptions, permissions, and review |
| Tableau Agent | Tableau web authoring | analysts building views and calculations in an existing workbook | works through the worksheet and data source, retaining a conventional authoring surface for correction |
| Looker Conversational Analytics | natural-language questions over LookML | business users querying governed fields and measures | administrators own glossaries, defaults, verified queries, and semantic-model quality |
| ThoughtSpot Spotter | search and conversational analytics | business teams asking everyday or strategic questions | provider positions data teams as maintainers of governed models rather than answer queues |
| Ask Sigma | conversational analysis over a cloud data workspace | non-specialists and analysts exploring organization data | shows data sources, formulas, filters, and analysis steps and allows step-level edits instead of full regeneration |
| Qlik Answers | structured and unstructured enterprise knowledge | business users asking questions and authors assembling analytic output | combines RAG, chart and dashboard agents, permissions, and audit logs; evidence remains provider-authored |
| Databricks Genie | a curated, no-code data chat | business users asking questions beyond existing dashboards | returns SQL-backed answers and visualizations; authors monitor, review, and refine instructions and definitions |
| Amazon Q in QuickSight | BI authoring, Q&A, executive summaries, and data stories | authors and consumers already in the AWS BI surface | integration reduces handoff; correctness still inherits datasets, topics, permissions, and author review |
| Oracle Analytics AI Assistant | natural-language analytics in Oracle | authors who prepare metadata and consumers who request charts or narratives | explicitly separates the author who prepares data and metadata from the consumer who asks questions |
Available controls are not governed deployment
Current product documentation establishes meaningful control surfaces. It does not establish how a particular organization configured, approved, monitored, and rechecked a particular deployment. A bounded seven-row audit found two real halves and no complete join:
- Power BI/Fabric, Tableau, and Looker document scoped enablement, data or semantic permissions, processing or masking choices, and some audit, feedback, or usage-monitoring surface.
- Provider-published stories report named-feature use of Looker Conversational Analytics at Google Cloud Support and Copilot in Power BI among KPMG developers.
- Two longitudinal studies add adjacent organizational-governance process evidence: a 2021–2025 HealthTech dashboard program and a 12-month ethics-based AI audit at AstraZeneca.
- Zero of the seven rows joins the exact AI-assisted visualization feature and configuration to authorization, monitored routine use, incident or exception disposition, an independently accepted outcome, and a later recheck.
The product-control evidence is concrete. Microsoft documents tenant, capacity, workspace, security-group, regional-processing, and Purview audit surfaces for Copilot in Fabric and Power BI; it separately warns that tenant settings help establish governance policy but are not security measures (enablement, Purview, tenant settings). Tableau documents site- and group-level enablement, data protections, masking, and optional audit-and-feedback collection (Cloud setup, Tableau Next setup). Looker documents instance, feature, model, data, content, and agent permissions, LookML grounding, usage monitoring, and regional processing; it also warns that answers can be plausible and wrong and tells customers to consult their authorizing body (setup, overview).
The organizational-use stories are also real but narrower. The Google Cloud Support story reports routine use over a governed semantic layer and claimed operational time savings, but does not expose the exact configuration, audit review, incident history, independent acceptance method, or later recheck. The KPMG story reports use of Copilot for narratives and prebuilt questions without a user denominator, exact configuration, feature-specific governance record, or independently attributed outcome. Neither story licenses a prevalence estimate or a provider comparison.
| Same-deployment receipt | What must be inspectable |
|---|---|
| Feature and version | Exact assistant, model or release coordinate, and date |
| Organizational authorization | Named approver, policy basis, and review date |
| Enablement and scope | Effective tenant, site, capacity, group, workspace, and task boundary |
| Data or semantic authority | Identity, permissions, governed objects, exclusions, and query route |
| Monitoring and audit | Events retained, reviewer, cadence, and escalation threshold |
| Routine use | Successful tasks and failures with an eligible-use denominator |
| Incident or exception disposition | What failed, who owned it, correction, and changed gate |
| Accepted outcome | Decision owner and independently checked operational consequence |
| Later recheck | Same lineage retested after material data, model, policy, or workflow change |
These receipts must meet on one deployment row. Product controls from one row cannot be combined with a different customer’s reported outcome to manufacture a governance score. The zero is bounded to the seven held public rows, not a claim that the products are ungovernable or that stronger private records do not exist.
Mixed-initiative and open-source systems
| Project | Approach | Why it matters | Evidence boundary |
|---|---|---|---|
| Data Formulator | direct encoding controls plus natural-language transformation, branching, visible derived data, and report sharing | treats AI as one control surface in an inspectable analysis environment | research evaluations are small; the current repository is more capable than the studied GPT-3.5 system |
| Lumen | declarative data pipelines, visualizations, dashboards, and specialist agents that can be serialized and edited | makes generated transformations and output inspectable and reusable across chat, notebooks, and dashboards | open-source feature surface, not independent evidence of routine field success |
| Vizro | low-code Python dashboard specification with code escape hatches and agent-facing tools | constrains generation to a production-oriented component system while preserving extension and deployment paths | a reusable framework reduces some errors; acceptance still requires tests and design judgment |
| AntV chart visualization skills | retrievable, version-specific instructions and declarative chart or code generators for several visualization libraries | shows how a harness can supply library syntax, chart vocabulary, and guardrails at generation time | the project reports its own 174-case testing; independent and current-model ablations are still needed |
Six techniques now cut across product classes: unrestricted code generation; semantic grounding in governed fields and measures; structured or declarative intermediate representations; direct manipulation and local undo; deterministic render or data checks; and proactive reader scaffolding. These techniques can coexist. The meaningful choice is which one contains the most expensive failure for the actual job.
A practitioner may therefore ask a general model for dashboard concepts, use a coding agent to generate measures, edit the layout manually in BI software, and let a reader-facing assistant explain the final report. That is one workflow with four evidence and control boundaries—not one assistant experience.
What feels amazing to creators
A first draft appears before the idea has cooled
The most consistent delight is immediacy. A person can move from a question or rough picture to something concrete enough to react to. In the DashChat study, participants valued rapid multi-chart mockups and simulated data as a way to negotiate dashboard requirements before production data existed. Data Formulator participants could combine direct encoding choices with short transformation prompts and complete all 16 reproduction tasks, although the eight-person study did not include a conventional-tool baseline.
This is valuable because a draft changes the conversation. “Show me three ways to compare this” is easier to critique than an abstract discussion of a future dashboard. The benefit is not that the first output is right; it is that the cost of making a candidate has fallen.
Tedious, legible work compresses well
Practitioners repeatedly describe value in SQL, calculations, documentation, bulk renaming, standard chart setup, tooltips, formatting, and alternative approaches. These tasks have inspectable outputs and a relatively short path to correction. In one public BI discussion, contributors described using AI to explain dashboard queries in documentation, prepare stakeholder questions, review junior work, critique mockups, and accelerate unfamiliar API or Python tasks. This is public testimony, not a survey, but it gives the broad phrase “AI in BI” concrete content.
AI can bridge representations a person does not fluently command
The 22-analyst verification study found that people moved among natural-language explanations, code, original data, intermediate data, results, and visual summaries. When Python was unfamiliar, some participants translated the logic into SQL or operations they knew. AI is useful here as a representation bridge: it can express the same intent as prose, code, a table, or a visual candidate. The person still needs at least one representation in which they can evaluate the result.
A bounded integrated tool can be genuinely faster
In a randomized public-health analysis exercise with 30 analyzed participants, the integrated ChatGPT environment finished faster than the distributed R/Stata-plus-ChatGPT workflow: a median 38 versus 45 minutes. The integrated group’s visualizations were also judged more meaningful and correctly labeled. That is real evidence for reducing tool-switching and syntax work in this specific 45-minute exercise. It is not evidence that the integrated group’s overall work was safer or better; that result points in the other direction.
A critic can improve work without becoming the designer
Visualizationary followed 13 designers—six novice, four intermediate, and three expert—across a three-to-five-day window as they iterated at least five versions of a visualization made with their own data and preferred tool. Its GPT-3.5 critique was combined with deterministic perceptual filters, a hierarchical report, and version tracking. Three external experts later rated the final work as improved on average (3.69 on a five-point scale where three meant no change).
The system did not supply a universally correct edit list. Participants used 68 pieces of feedback, ignored advice that conflicted with intent or seemed wrong, and did not always end with their best-rated version. Intermediate and expert designers generally converted critique into better revisions more effectively than novices. The small, short study has no conventional-critique baseline, but it demonstrates a credible role for AI as inspectable counsel inside an existing practice rather than as an autonomous author.
What makes the work hard
The model can make an invisible analytical decision
Aggregation, filtering, grouping, baseline, and scale choices can be buried in generated code or made implicitly from an ambiguous prompt. In the 2026 novice study, 13 of 20 participants issued ambiguous or problematic prompts, and the system sometimes chose inconsistent aggregations or supplied interpretations the user could not verify. A chart can look conventional while encoding a different question from the one its creator thinks it answers.
Governed BI environments try to narrow this gap through semantic models, descriptions, verified queries, permissions, and authored instructions. The current Looker documentation is unusually explicit: the model maps natural language onto maintained LookML fields and lets administrators supply business glossaries, defaults, and verified queries. That is an architecture contract, not proof of perfect answers. It also makes the organizational dependency visible: conversational analytics inherits the quality of the semantic model someone maintains.
Correction can be harder than creation
Natural language is a powerful coarse control and often a poor fine control. In the novice study, 11 of 16 attempts to fix clutter and eight of nine attempts to repair unusable charts failed. More capable systems generated richer interfaces, but also introduced longer waits, unclear selected states, broken controls, and cases where a model interpreted a blank rendering as if it contained data.
Public Power BI accounts add a practical version of the same problem. In a 35-entry thread, contributors described useful greenfield structure alongside stale filters, bookmark IDs, whitespace cleanup, slow preview cycles, and minor changes that were faster to make manually. One contributor summarized the division as trusting and verifying generated code while retaining human judgment over whether business users would accept the report. These are cases, not frequencies.
Another practitioner reported a generated inline calculation that produced a desired bar chart but could not be located or edited through familiar Power BI mechanisms. The issue is not merely whether the number happened to be right. The generated implementation broke the user’s ability to inspect, maintain, and explain it.
The longitudinal Visualizationary study found a quieter form of the same cost: feedback could be verbose, vague, or based on a mistaken reading of intent. Novices in particular struggled to translate broad advice into a precise edit, and the final version was not always the highest-rated one. A critique system therefore needs triage, rationale, and edit locality; more feedback is not the same as a better correction path.
A 2024 controlled study of AI-assisted data analysis gives this problem an actual, if deliberately artificial, denominator. Eighteen experienced analysts completed 108 task episodes designed so that the model would make two errors unless the person found and repaired them. Seven episodes were not completed within 15 minutes. In 31 episodes—28.7% of the study tasks—the participant said the work was complete while an issue remained. Interfaces that exposed editable assumptions and plans improved participants’ sense of control, but the study did not detect a difference among conditions in task success, completion time, or verification hints.
This is evidence that premature acceptance and non-completion occur even among experienced analysts who have been warned to look for errors. It is not a field failure rate: every task was engineered to be troublesome, the researcher supplied correctness checks and hints, and no episode continued through publication or abandonment in ordinary work.
Total cost includes waiting, prompting, cleanup, and model usage
First-pass generation time omits the rest of the loop. Practitioners mention prompt construction, regeneration latency, application refreshes, token or credit consumption, deployment cycles, and the cognitive cost of deciding whether a strange result is a model error or their own misunderstanding. A one-minute generation can be slower than a ten-second direct edit; a 25-minute dashboard rebuild can be impressive and still depend on an existing screenshot, prepared tables, and substantial layout review.
The useful unit is therefore total human-plus-system time to an accepted artifact, including correction and verification—not seconds to first render.
That unit needs two receipts. The numerator is all reconciled route cost over the declared work: successful attempts, rejected drafts, abandonment, quarantine, and no-output cases, plus the declared share of preparation and ownership. The denominator counts only artifacts that passed the frozen human or production acceptance contract. Tool calls, a per-chart API average, aggregate benchmark accuracy, candidate production, or an agent’s finish signal supplies neither receipt. The specialist-cost register now holds eleven partial fragments and zero accepted-artifact cost comparisons. For a practitioner, the practical question is not “how cheap was the average call?” but “what did every attempt cost, and how many became work I could accept?”
One fragment clears a narrower budget bar. The peer-reviewed Selective TTS visual-insights study matches declared LLM calls and, separately, output tokens across pruning policies in one pipeline. Its released implementation still leaves code-repair calls and errored-worker cost outside aggregate totals, records completion tokens only, and counts proxy-judge-scored candidates rather than contract-accepted work. Practitioners therefore need three separate receipts: the promised partial budget, what every attempt actually used, and what accepted work cost. Zero of eleven held fragments supplies the last two for every compared arm.
The same 108-episode study illustrates why interface improvements should not be translated automatically into time savings. Mean completion time was 543 seconds for the conversational baseline, 588 seconds for stepwise decomposition, and 658 seconds for phasewise decomposition; the differences were not statistically significant. More visible structure made correction feel more controllable while also creating intervention and information costs. The study measured human task time, not model inference, subscription cost, later cleanup, or maintenance, so it still does not provide a total-cost ledger.
Production adds failure surfaces that studies often remove
Controlled studies usually provide clean data, bounded tasks, and a working prototype. Real work adds permissions, confidential data, incomplete semantics, unreliable connectors, browser behavior, responsive layout, organizational review, publishing systems, and future maintenance. The 2025 randomized exercise recorded upload and browser problems even in a simulated setting. Practitioners in public threads describe network policy, application refresh, deployment, and cost constraints that are normally absent from benchmark scores.
Interviews with 17 biomedical-visualization practitioners show where a high-stakes production boundary is currently being drawn. Thirteen already used generative AI somewhere in their workflow, primarily for auxiliary work such as research, moodboards, boilerplate code, background assets, captions, or translation. All opposed incorporating substantial AI-generated visual content into final scientific work, citing validation, accuracy, copyright, provenance, and accountability. That is evidence about real production and dissemination judgment, not delivery efficacy: the study did not observe projects through handoff, publication, reader use, refresh, or maintenance.
Speed and quality separate in the measured studies
The most cautionary evidence is not that systems always fail. It is that their success signals can disagree.
| Study | What improved or looked good | What did not follow |
|---|---|---|
| Randomized public-health analysis exercise, 2025 — 30 analyzed participants | Integrated AI users finished seven minutes faster at the median; their charts were judged more meaningful and correctly labeled. | Overall scores were not significantly different. Only 6.7% of integrated submissions versus 26% of distributed-tool submissions were free of serious errors. Small, underpowered study; no non-AI control. |
| Vibe Visualizing, 2026 — 20 novices, 60 sessions, 175 charts | Participants could repeatedly produce charts and reported mean confidence of 3.73/5 and satisfaction of 3.93/5. Replayed prompts showed better visual output from newer models on some measures. | Every chart had at least one design flaw; 52 of 60 sessions had fatal task noncompliance; 22 charts were unusable; incorrect insights appeared in 12 sessions. Verification was rare and many repair attempts failed. |
| Data Formulator 2, CHI 2025 — eight corporate participants | All participants completed 16 reproduction charts; direct controls, short prompts, visible transformed data, history, and multiple inspection artifacts supported work. | Six of eight still needed hints. The tasks were reproduction, not open analysis; no conventional-tool baseline or self-owned data; trust in simple steps could carry too easily into complex descendants. |
| DashChat, 2025 preprint — formative interviews plus 28-person evaluation | Rapid pre-data dashboard prototypes helped make requirements and layouts discussable; designers rated it easier and more effective than a lightly taught Tableau condition. | The comparison favored the familiar purpose-built workflow and measured mockups, not real-data correctness, deployment, or ongoing dashboard use. Precise edits and domain conventions remained difficult. |
These results should not be pooled into a universal score. They agree on a more useful point: assistance can improve access, speed, or perceived ease without proportionately improving semantic correctness, error freedom, or delivery.
Expertise is both leverage and burden
The simple story that novices benefit and experts do not is wrong. So is the story that experts automatically get the best results.
Visualization expertise is not one rank. The strongest current literacy framework separates the ability to consume, construct, critique, and connect a visualization to its context. Real work also draws separately on data and statistical knowledge, domain semantics, visual design, implementation, situated judgment, and delivery experience. A person can be a strong business reader and a weak programmer, a domain expert and a novice chart author, or an expert implementer who does not know the local measure definitions. AI removes different friction for each profile.
The human-skills companion therefore uses a stricter definition of a gain. Access, productivity, artifact quality, independent learning, verification, reader outcome, and sustained work are separate. Confidence or an AI-assisted final project does not establish learning; time to a first render does not establish productivity to an accepted artifact.
The learning companion follows that distinction across time. It finds that implementation, translation, and on-demand explanation are becoming easier, while delayed independent construction remains mostly unmeasured. It recommends removing inspectable production friction while preserving prediction, data-to-encoding reasoning, comparison, diagnosis, local repair, retrieval, and unassisted transfer.
Experienced practitioners can filter and redirect
In the comparison between ChatGPT advice and human visualization experts, practitioners valued AI for rapid brainstorming, broad option lists, and a neutral first pass. One experienced participant’s description was roughly: ask for 20 ideas, discard 16, use four. That workflow is valuable because the person can perform the filtering. In the analyst verification study, people with data experience inspected distributions and values, while coders inspected code and others translated logic into familiar operations.
The longitudinal designer study adds direct evidence. Intermediate and expert participants were better able to select among critiques and enact useful changes, while novices needed more help turning high-level feedback into design operations. A 2025 single-company case study reached a compatible but much weaker result: eight interviews produced basic, intermediate, and advanced profiles, then one person from each profile used a prototype for about 30 minutes. The basic user needed initial prompting support and missed some interaction affordances; the advanced user demanded source transparency, encountered a wrong scatterplot, and formulated more intricate corrections. Three probe sessions cannot establish efficacy, but they show why one interface cannot equate “easier” with less visible state for every user.
The same expertise exposes prompt overhead
Human experts were still preferred to ChatGPT for accuracy, helpfulness, reliability, adaptability, context, and actionable advice in the 12-practitioner study. The model produced broad, agreeable recommendations; experts narrowed the scope, built common ground, and raised issues the practitioner had not known to ask about. The model sessions were conducted in 2023, so their raw quality comparison is dated. The contextual lesson remains visible in newer public accounts: experienced people often reserve AI for repetitive or unfamiliar work because describing a familiar small edit can cost more than doing it.
Novices gain access without gaining an error detector
The 2026 novice study is the clearest warning. Participants could produce plausible charts, but verification intent was low, only three verification attempts were observed, and satisfaction remained positive. Some users followed inappropriate model suggestions. Poor visualization and data literacy affected both prompting and the ability to diagnose the result.
This creates an expertise paradox: the people for whom the interface removes the largest access barrier may be the least able to tell when it has quietly answered the wrong question. Better defaults help, but they do not eliminate the need to expose the decision the default made.
The reader receives a claim, not a generation process
Creators experience prompts, waiting, code, and corrections. Readers usually see only the final artifact in a report, article, dashboard, presentation, or social post. They infer care and authority from what is visible.
Trust is assembled from cues
In a 2025 preprint, 37 US participants ranked static charts from news, science, government, and infographic contexts and explained their choices. Clarity was mentioned by 31 of 37 participants; source citation, familiar chart forms, integrity, and visual polish also mattered. Priorities were consistent within many individuals and different across individuals. Infographics polarized viewers: some saw accessibility, others saw promotion or clutter.
The study measured deliberative, self-reported trust—not truth detection, comprehension, or behavior. That distinction is decisive for AI-generated charts. A clean chart can signal professional care even when no one inspected its transformations. A cited source can increase confidence without proving that the chart represented it faithfully.
An AI label does not have one predictable effect
The preregistered 2021 Vis Ex Machina experiment showed identical chart panels with human or algorithm labels to 114 participants. Before the task, 60% preferred human recommendations and 15% preferred algorithmic recommendations; actual choices were approximately even. Data relevance dominated, but some participants followed their prior source preference. An algorithmic label suggested precision or error-freedom to some people and lack of human judgment to others.
That study predates generative visualization and does not tell us whether to label current charts “AI-generated.” It establishes that provenance cues interact with prior beliefs and task context; disclosure is not a universal trust switch.
A model reader is not a human-reader test
In a 2026 study of 60 synthetic charts, three multimodal models matched a narrow designer-intent label more often than 24 human participants. The models tended to enumerate structure and values; people formed trend narratives and were more affected by layout and overlap. The finding is not that models are better chart readers. It is that the metric rewarded a task aligned with machine decoding while human reading pursued a different form of meaning.
Using a model to inspect a chart can catch useful defects. It cannot establish that a journalist’s audience understood the story, a dashboard user noticed an alert, or a student learned the concept.
A human study still needs the AI role and artifact lineage
A bounded audit of ten held primary rows separates evidence that is often collapsed under “human evaluation.” Only ChartAttack carries selected AI-generated chart instances into an independent controlled-reader outcome. Five other rows measure real human trust, comprehension, decision, or access outcomes, but AI labels, explains, assists, or sits beside an existing visualization. Three rows follow creators or co-designers without an independent audience outcome. One technical evaluation uses VQA and no human reader.
The complete chain therefore needs eight receipts on one artifact lineage: AI role; exact artifact, build, and model; acceptance and delivery; independent intended reader; intended context and access path; measured outcome; attribution or comparison; and later recheck. Across the ten held rows, one direct controlled reader effect exists, but zero binds an immutable artifact and model to accepted delivery, representative intended-context use, and a later same-lineage recheck. A creator interpreting their own chart, a blind participant using an AI explanation over an existing chart, and a VQA model scoring a generated chart are three different evidence states.
Longitudinal evaluation can inform deployment without proving audience consequence
Lexara adds a stronger middle receipt. Six conversational visual analytics developers used a deployed evaluation toolkit for two weeks, running 38 experiments over 57 uniquely authored test cases, ten models, and six system prompts. They used their own data and prompts and logged model and prompt selections, rationale, outputs, observations, and confidence. One participant switched models after a JSON diff exposed an encoding mismatch; another challenged the toolkit’s recommendation.
That is longitudinal development-decision evidence. It is not yet a consequential audience decision or calibrated trust. The study does not bind an immutable participant-tested toolkit build and exact per-experiment configuration to a contract-accepted visualization, actual downstream delivery, intended-reader outcome, whole cost, and later post-release recheck. Its authors describe the metrics as diagnostic signals rather than proxies for human sensemaking, and leave cost, latency, and prompt/model drift outside the study.
The combined decision/recheck audit keeps twelve receipts on one lineage: AI visualization role; immutable artifact and build; exact model and configuration; accepted artifact; intended-reader delivery; context and access path; consequential decision; calibrated trust; elapsed later event; same-lineage recheck; actor-separated maintenance authority; and whole cost. Across the ten reader-outcome cases, three production afterlives, Lexara, and the later same-feature InfoViP operational line, there are 15 named cases and zero full twelve-receipt episodes. The useful next move is to carry a Lexara-style evaluation export into a frozen release and return-user check—not to borrow ChartAttack’s reader effect or a different dashboard’s production afterlife.
A pin-able output bundle is still not audience evidence
The first bounded screen inside the review-owned frame makes one useful artifact-level upgrade without completing the lifecycle. Of A208’s 122 authoritative stable-key titles, one explicitly signals a generative-AI study: Beyond Generating Code: Evaluating GPT on a Data Visualization Course. The paper repeats 91 quiz questions ten times, tests nine homework assignments with GPT-3.5 and GPT-4 aliases, and uses three former teaching fellows for 54 homework-grade assessments. GPT-4 reaches 86.4% on quizzes and 79.7% on homework; graders correctly identify 38 of the 54 submissions, misidentify eight, and are unsure on eight.
The linked supplemental-output
repository
is materially better custody than a results paragraph alone. At the 16 August
2026 evidence cut it exposes one pinned root commit containing 1,271 blobs,
including quiz prompts/results, homework prompts, JSON output, generated code,
and grader-ready bundles. The quiz runner records a gpt-3.5-turbo alias and a
commented gpt-4 alternative, but not an immutable provider model snapshot or
complete per-call configuration. The repository exposes no named release or
tag and no later commit against which to recheck the bundle.
The paper also stops before delivery. It does not directly evaluate the course’s final project. Prompts ask GPT to write for governments, the United Nations, and primary-school students, but none of those intended audiences receives or evaluates the resulting artifact. Retries are described—including up to 30 retries for browser errors—without whole-workflow cost. The result is therefore one pin-able output bundle and zero audience-delivery or later- recheck chains, not a deployed-audience outcome.
This was a title-cue screen, not a full-text exclusion pass. The other 121 stable-key titles did not surface an equally explicit GenAI cue; they are not 121 negative findings. Five supplement positions also remain outside the authoritative frame. The complete review-lifecycle denominator stays null.
A mature visualization afterlife still cannot supply the AI role
The next DOI-level discovery pass deepens coverage without completing the review denominator. Of A208’s 122 authoritative stable keys, 114 carry DOIs. Exact-DOI OpenAlex retrieval returns all 114 records; 87 expose abstract indexes and 27 do not. The abstract cue screen adds no hidden GenAI candidate beyond the already known GPT course paper. Its 86 nonmatches are not surfaced, not excluded. Eight stable non-DOI keys remain outside this batch.
One available abstract does surface an unusually mature delivery record: AIDSVu, a public HIV-data mapping and dissemination platform launched in 2009. Its primary paper reports ten years of evolution, 501,527 unique website users in 2019, nearly four million social impressions, 249 peer-reviewed publications using or citing the resource, named government/nonprofit/academic audience proxies, and uses in testing campaigns, telemedicine siting, service-gap identification, grants, and research-site selection. Governance groups, annual data commitments, privacy suppression, and a maintenance-funding boundary are explicit.
The afterlife is current. AIDSVu’s official 24 June 2026 data release publishes 2025 PrEP data through interactive maps, profiles, and PrEPVu.org; its current FAQ distinguishes unavailable, unreleased, suppressed, lagged, estimated, and later-corrected data and commits to ongoing and annual updates.
These are stronger public delivery, use, governance, and later-event receipts than most AI cases in this report. They still expose no AI visualization role, immutable artifact/build, version-bound measured reader or decision outcome, calibrated trust, affected-audience post-change recheck, or whole cost. Aggregate reach, IP-derived organization types, and self-reported use examples are not one actor’s accepted outcome. Retain AIDSVu as a non-AI delivery/afterlife comparator; do not make it a fifteenth AI-assisted case or borrow its mature platform history into a model feature.
A human study, co-design, and production intent are not deployment
The next primary-text pass resolves part of the 27-row no-abstract layer. Two repository-full-text signals and three additional exact-title captures put five full texts into custody: four empirical papers and one review manuscript. All five are non-GenAI on the explicit GPT, ChatGPT, generative-AI, large- language-model, and LLM cue screen. Twenty-two DOI rows and eight stable non- DOI keys remain unassessed, not negative.
The useful result is not one composite success story. It is a receipt ladder:
| Receipt lane | What the primary text establishes | What it does not establish |
|---|---|---|
| Controlled audience measurement | B17 measures accuracy, time, NASA-TLX, and working-memory load with 23 participants. B128 measures task time, errors, and cognitive effort with 146 decision-task participants and 107 questionnaire respondents. | A released product, accepted organizational decision, consequential field outcome, or later audience return. |
| Public-sector co-design, demo, and bounded use | B120 records three months of HHS co-design, named officials, a final demo, and a short task session. An archived White House plan independently records prototype version 1.0 and named audience access. | A defined usability denominator, production acceptance, scalability, security, sustained updating, accessibility, or decision impact. The official plan says the prototype does not replace professional-grade software or services. |
| Enterprise demonstrator and production intent | B90 records a multinational-company test bed, explainable clustering, human rule validation, an implemented demonstrator, and initial profile-group evaluation. The partner was working toward commercial production. | A participant denominator, exact released build, production acceptance, routine use, outcome, later event, affected-user recheck, or whole cost. |
The fifth text is a review of open-government-data visualization, not a fifth artifact episode. It finds heterogeneous evaluation and only three formal usability evaluations among eight tool/framework studies. That context reinforces the ladder but cannot lend one reviewed artifact’s evidence to another.
Co-design, a demo, a human task study, and production intent are four different receipts—not deployment. Keep them attached to their exact artifact and stage. A subsequent high-signal recovery adds eight abstract or official-project content surfaces and reduces the substantive-content gap from 22 DOI rows to 14; it does not add a deployment. A deployment claim still needs an immutable accepted build, delivery to named users, a consequential outcome, a later material event, and a return by the affected audience.
A field-use receipt is still not an AI lifecycle
The eight stable keys outside the DOI-bearing screen were partly an authority- repair problem, not an evidence void. Exact publisher and identifier records recover three DOIs, correct B136 from 2014 to 2012 and B143 from 2021 to 2020, and show that the PMID printed for B13 names an unrelated medical-physics abstract. Four full primary texts, two primary abstracts, one publisher-issue excerpt, and one gated primary text are now distinguished. The gated row remains unassessed rather than negative.
The 2012 cancer diary is the strongest practice receipt in this stratum. Its primary abstract says system logs and a survey demonstrate increasing use and good usability for a summarized oncology-history view inside Siemens Soarian Clinicals EMR. An independent 2025 systematic review recovers the human- evaluation details: 11 clinicians, a five-point Questionnaire for User Interaction Satisfaction, and median overall satisfaction of 4.38.
That is more mature than a concept, staged task, or partner demonstrator. It still supplies none of the receipts needed to call it an AI lifecycle: no AI role, immutable accepted diary build, system-log base, patient or decision outcome, later material release or incident, affected-clinician return, or whole cost. The other held rows preserve separate lanes—one bounded emergency- operator simulation, two technical experiments/demonstrators, and three conceptual or unevaluated prototypes. Their receipts cannot be combined.
Seven held content surfaces contain no explicit GPT, ChatGPT, generative-AI, large-language-model, or LLM cue. That bounded null does not classify the gated eighth row, all 122 stable keys, or all 127 supplement positions. Keep the cancer diary as a non-AI embedded-use comparator, not a complete lifecycle.
Seven human evaluations are still not one deployed lifecycle
Nine of the 22 previously uncaptured DOI rows were selected because their titles explicitly signal users, trust, perceptions, decision support, or evaluation. Eight now have substantive primary abstracts or official project records in custody, seven explicitly report participant or evaluator studies, and one course-choice dashboard adds a formative pilot whose direct dashboard exposure is ambiguous. One forest-planning row remains metadata-only and unassessed. The substantive-content remainder falls from 22 to 14; none of the new surfaces is a full paper.
The studies answer different questions. A 22-participant IVR case reports bounded bridge-damage identification and acceptable usability. A three-stage Finland-China experiment tests trust and reputation indicators against willingness and checking. Eleven public health nurses evaluate patient-level prototypes. Other rows use student eye tracking, real fantasy- football users, or chart-comprehension measures. These are genuine human receipts, but they are controlled tasks, prototype evaluations, or formative pilots—not routine operational use.
InfoViP is the strongest delivery near-miss. Seven FDA safety evaluators supplied requirements, reviewed wireframes, evaluated the prototype, and had their suggestions addressed. The primary abstract identifies natural-language processing; an official FDA project page also describes unsupervised-learning work and says an enhanced version will be installed in production. The future tense matters. The held record does not confirm installation, authorization, accepted build, routine use, a drug- safety or regulatory decision outcome, later change, or evaluator return.
No held content surface in this batch contains an explicit GPT, ChatGPT,
generative-AI, large-language-model, or LLM cue. InfoViP is an AI-adjacent,
non-generative evaluator/prototype/production-intent comparator. Keep
evaluated, planned, installed, accepted, used, outcome-bearing, and
rechecked as separate states. Seven human-evaluation signals are still zero
complete production lifecycles.
An operational data pipeline still needs accepted implementation and user return
Later records materially advance the same InfoViP feature beyond the earlier production-intent statement. A full 2022 publisher article documents a standalone deduplication web program deployed on FDA servers for a voluntary six-month evaluation in reviewers’ actual day-to-day work. About 60 reviewers were invited; 20 unique reviewers made 58 submissions; 27 feedback files returned, 26 with labeled duplicate determinations. Mean pairwise recall was 0.71 (SD 0.32) and precision 0.67 (SD 0.34). The five largest submissions, each above 1,900 reports, returned no feedback. This is a counted real-work field evaluation, not a current routine-use denominator or post-change return.
A later 2025 indexed primary abstract says the upgraded deduplication pipeline was applied to 29 million historic reports and daily incoming reports, identified about five million duplicates, and is operating at FDA. Across 12 human-expert-adjudicated datasets totaling 2,300 reports, it outperformed the current tool on ten; F1 ranged from 0.36 to 0.93, with half the sets above 0.75. The full paper is not held, so this operating and performance evidence remains abstract-bounded. It establishes component operation, not acceptance or authorization of every InfoViP interface and feature.
Two intervening public reports remain part of the same chronology. FDA’s 2022 OSE annual report calls InfoViP an operational pilot with temporal visualization, reviewer- confirmed NLP duplicate detection, and an NLP/ML assessability classifier. A 2024 conference abstract reports assessability F1 above 0.85 and human quality-control affordances, but does not count its reviewer-confidence claim or identify a tested build and consequential decision outcome.
Three official FDA award rows bound one external cost envelope for the named
line: FY2019 contract
75F40119C10084 lists
$1,026,822 obligated and estimated; FY2021
75F40121C00185 lists
$926,349; and FY2023
75F40123C00113 lists
$447,337 obligated against $911,251 estimated. Together they total $2,400,508
obligated and $2,864,422 estimated. They omit FDA labor and
infrastructure, grant support, reviewer correction, operations, maintenance,
and failure cost. They are not whole cost and provide no ROI denominator.
The current HHS AI use-case
inventory
supplies the counterevidence. It records Acquisition and/or Development,
Date Implemented: N/A, no associated AI-system Authority to Operate,
partially completed training/evaluation-data documentation, and limited
internal review documentation. Its note that InfoViP reuses production-level
code from another use case is not proof that InfoViP itself is authorized or
implemented in production.
This is the first held AI-and-visualization line that joins practitioner co- design, counted internal real-work field evaluation, later operating-scale processing, comparative expert-adjudicated performance, a public award envelope, and current governance counterevidence. It remains the fifteenth named practice or afterlife case. It still has no immutable accepted whole- platform build, current routine post-release user denominator, calibrated trust result, attributable regulatory outcome, later affected-reviewer return, or whole human-plus-machine cost. Do not attribute organization-wide FDA outcomes to InfoViP; no held source makes that causal join.
The final report supports a production component and exposes its QA afterlife
The final 2025 CIOMS Working Group XIV report crosses a stronger, component-level threshold than the earlier indexed abstract. Its case-deduplication use case says the pipeline was approved for historical and incoming-live ETL processing, installed in AWS, tightly integrated with the agency adverse-event reporting system, and by 30 July 2025 had screened more than 30 million historical reports while continuing to deduplicate about 8,000 new submissions daily. It says case-series use is fully implemented, medical experts can confirm or modify the reference case in that routine workflow, and administrators continuously monitor the pipeline and use of its output.
Keep that promotion bounded. The same final report says the complete daily
output cannot be human-confirmed, a solid QA plan was not yet in place,
periodic audits were planned rather than reported as completed, routine-use
roles were not fully assigned, and effects on data-mining calculations and
safety-signal discovery had not been investigated. It says underlying code is
available to the regulator; no public immutable build, formal acceptance or
post-change sign-off artifact, or maintenance handoff is held. The current
HHS inventory
still records implementation N/A and no AI-system ATO. The records may differ
by component, broader use case, or date; do not invent a reconciliation.
An official 2026 FDA fellowship record proposes integrating Elsa GenAI capabilities through APIs and building an enhanced InfoViP interface. Its flexible anticipated start and two-year appointment are prospective. No held receipt says a fellow started or that the integration was built, released, exposed to reviewers, or used. Treat any such release as a named change requiring fresh version, authorization, QA, ordinary- use, and affected-reviewer evidence.
The surrounding dates and roles need one explicit repair. An FDA CDER Conversation appears 2026-dated in one HTML field, but the visible “Content current as of” field and modified date both say 3 April 2024. Its statement that InfoViP “will allow” safety reviewers to focus on complex data belongs before the final 2025 report; it is not a 2026 regression. A May 2026 FDA speaker biography names Oanh Dang as serving as InfoViP project lead. Keep that precise role separate from service operator, maintenance owner, release approver, and authorization authority. FDA’s Elsa 4.0/HALO announcement confirms an agency-wide Elsa release and that HALO integration began, but does not name InfoViP. Broad platform availability is not an application-specific build or exposure receipt.
The live fellowship page now adds an administrative afterlife: the opportunity is closed after its 8 May deadline. It still reports no selected fellow, onboarding, start, build, release, exposure, or use. It names Joshua Xu, Leihong Wu, and Oanh Dang as mentors for questions about the nature of the research. FDA coordinates Xu to Research-to-Review and Return and advanced-AI integration, Wu to AI/ML bioinformatics research, and Dang to InfoViP project leadership. These are useful research contacts, not a service operator, maintenance owner, QA executor, release approver, or authorization authority. The appointment terms reinforce the boundary: a fellow would be a nonemployee prohibited from inherently governmental functions. An overlapping FDA-CERSI QA project is still described as a literature review and planned technical report; its page never names InfoViP or supplies a completed audit.
The defensible classification is one self-reported approved and production- integrated component with a bounded routine case-series path and administrator monitoring. The audit remains 15 named cases and zero complete episodes because immutable version, formal sign-off, ATO, completed QA audit, current user denominator, attributable outcome, affected-user return, and whole cost remain unjoined.
The residual closes to three records, not to a deployed lifecycle
A residual pass investigates every one of the 14 DOI rows left without substantive primary content. Eleven now have primary content in custody, including one full publisher chapter; B92, B91, and B110 are metadata-only at that evidence cut. The historical ledger is six full texts, 18 primary abstract or official-content rows, and three metadata-only rows. A later recheck changes B110 without rewriting this pass.
The clearest new mechanism is the open full chapter on a gesture-based data- center visualization. Three UX experts heuristically evaluated its second version and found 25 unique problems; the team used the findings to implement a third version. This is a strong evaluation-to-redesign receipt because it binds evaluator, problem inventory, and changed artifact state. The paper leaves real-user evaluation to future work. Its phrase “deployed in a large display” describes the display environment, not an accepted field deployment.
Two abstracts move closer to practice without crossing that boundary. A territorial-transformation case applies interactive spatial visualization with a multinational stakeholder group in the Genoa–Rotterdam corridor. A network-security case starts from university network-manager requirements and evaluates a developed tool in several university units. Neither surface identifies an accepted build, ordinary-use denominator, consequential corridor or security outcome, later material event, or affected-actor return.
Six of the 14 rows explicitly name users, experts, study participants,
stakeholders, or managers. One more reports a useful supplier-dashboard
evaluation without naming its actor or sample. The remaining substantive rows
cover pattern evaluation, example-data comparison, a clinical review, and an
optimization-backed system proposal. Simulated annealing is not generative AI,
and none of the 11 content surfaces explicitly assigns AI a visualization
role. Preserve the useful stages—expert-evaluated, redesigned, applied
with stakeholders, and evaluated across units—without shortening them to
deployed, adopted, or proven impact.
One interface is live; that is delivery, not use
The exact final-three recheck moves B110 out of metadata-only custody. Its official bibliographic abstract names stakeholders, policy makers, scientists, educators, resource managers, and users as intended audiences for DiscoverFramework, DiscoverWater, and DiscoverHABs. The framework source and DiscoverWater application are available at immutable commits. The named DiscoverWater KU interface returns instructions for a year slider, charts, and gauge-station exploration in August 2026.
That is a real same-lineage public-delivery afterlife. It is not evidence that the interface was continuously available or used from 2021 through 2026. No held manifest binds the current live bytes to either repository SHA. No source records owner acceptance, traffic or repeat-use denominators, a resource- management decision, learning or accessibility outcomes, a later affected- audience return, or whole creation and maintenance cost.
At that pass, B92 and B91 remained content-unassessed after exact publisher, proceedings, repository, API, researcher-network, and related-author routes failed to put their substantive text in custody. The 27-row ledger was then six full texts, 19 primary abstract or official-content rows, and two content-unassessed rows. A later exact publisher-abstract recovery moves B92 without changing the B110 delivery boundary below.
Partial source lineage still does not identify the deployed build
A deterministic follow-up sharpens that boundary. Reconstructing the only
published DiscoverFramework_v1.2 page produces a different digest and
materially different tour, Mapbox version, Syracuse layer, chart, and layout
code from the dated live page. Each published page component has only its
initial commit, the repository has one branch and no tags, and the separately
linked v2.0 repository is an R/Shiny application rather than the live
JavaScript/PHP page.
The result is not “unrelated source.” Eight of twelve local data-dependency basenames named by the live page occur in the pinned DiscoverWater tree. Of three live/repository asset pairs compared, one matches byte-for-byte, another matches after canonical JSON formatting, and one differs materially. That is partial same-lineage data reuse—not an exact deployed-build join.
The official 2018 prototype record describes a demonstration and author interpretation, while the current page contains analytics and comment hooks. Neither surface supplies a version-bound participant or traffic denominator, acceptance, consequential audience outcome, accessibility acceptance, later affected-audience return, whole cost, or AI visualization role. For reusable public work, preserve a deployment manifest, exact dependency digests, acceptance, telemetry, audience task, and later return as distinct receipts.
One publisher abstract closes one content gap; one retrieval loop stops
B92 now has lawful substantive primary content in the exact publisher record. Its abstract describes a multicriteria risk model for hydrogen pipelines, Monte Carlo simulations, Kendall’s tau comparisons, and graphs intended to help rank pipeline sections and target risk mitigation. That is a reusable non-AI decision-support method: propagate uncertainty, compare the stability of competing rankings, and show the decision order rather than only a score.
The exact abstract names no human participant or evaluator sample and assigns no role to GPT, ChatGPT, generative AI, an LLM, or AI-assisted visualization. It also supplies no accepted release, intended-audience delivery, observed-use denominator, consequential outcome, later recheck, or whole cost. Those are abstract-bounded absences; the full paper is not in custody. B91 remains content-unassessed because its exact publisher preview ends before the paper and no lawful substantive primary surface was recovered. Three bounded lawful passes now cover the exact DOI and retired CRC route, Crossref registration, publisher preview, current book, request-only researcher record, and related thesis. Repeating those routes is paused at the 16 August evidence cut.
The current 27-row ledger is therefore six full texts, 20 primary abstract
or official-content rows, and one content-unassessed row. Twenty-six rows can
now be screened at a substantive layer, and none exposes a complete AI
visualization lifecycle. Practitioners can reuse B92’s uncertainty-and-ranking
mechanism; analytics leaders still need acceptance and outcome receipts;
researchers should report 6 / 20 / 1 / 0; journalists should call it an
abstract-level non-AI comparator; builders need the implementation and tested
build; educators should distinguish method, user, delivery, and access; and
sponsors should fund B91 access only when an exact publisher abstract/body,
accepted manuscript, correction, or new lawful deposit can be obtained;
otherwise they should fund a complete same-artifact AI-assisted
delivery-to-recheck episode. The content ledger stays 6 / 20 / 1 / 0; its
separate workflow state is 1 access-blocked and unassessed / 0 active generic
B91 targets. Paused is not negative, and B91 stays in the denominator.
A review corpus is not a lifecycle denominator
A 2025 systematic review of data visualization in AI-assisted decision-making shows that the adjacent field is substantial. Its PRISMA search covered five academic databases through 1 July 2024 and retained 127 studies: 118 empirical papers and nine reviews. It maps visualization approaches, domain-expert challenges, visual elements, and evaluation methods.
It cannot establish how many studies complete the narrower lifecycle in this report. The authors explicitly say the review does not specifically focus on generative AI. Its article-level synthesis does not jointly encode a generated artifact and immutable version, model exposure, accepted delivery, representative decision-maker, consequential decision quality, calibrated trust, and later same-lineage recheck. The review names GenAI trust and interpretability as future investigation and longitudinal effects on decision behavior as future research.
The lifecycle count is therefore unknown, not zero of 127. The review is a useful coverage and discovery map. Turning fields it did not extract into negative results about all 127 primary papers would manufacture a denominator.
The review’s own three-page
supplement
makes that audit gate stricter. Its six domain totals add to 127, but Appendix
Table 2 names 126 author-year citations and prints one literal ? as the final
Education entry. A source-internal reconciliation now repairs the missing name
without rewriting the supplement. The review says domains were assigned after
selection and names five Education examples; four are already in the twelve-
name row, while Hernández-Calderón et al. (2023) alone is absent. The review’s
bibliography binds it to DOI
10.1093/iwc/iwac043, and Crossref
confirms that identity.
Keep both fields: published: ? and repair overlay: Hernández-Calderón et al.
(2023). The complete six-domain crosswalk then exposes the larger boundary.
The supplement’s 127 positions map to 122 distinct A208 review keys: 114 with
DOIs, two with confirmed PubMed identifiers, five without a publisher
identifier, and one with a printed PMID that resolves to another paper. There
are no duplicate authoritative keys.
Five supplement-only rows have exact-looking external DOI identities but no A208-controlled join: Jiao (2022), Burt et al. (2017), Theis et al. (2018), M. Lu (2020), and Islam et al. (2022). A registry can establish each candidate paper’s identity, not the review authors’ intent. The B13 PMID is a different failure mode: its numeric identifier resolves to an unrelated Medical Physics record and must not be reused.
Keep five fields distinct: source label, authoritative review key, external candidate, identifier validation, and admission authority. Obtain a review- author, publisher, extraction, or A208-controlled join for each of the five rows and repair or remove the conflicted identifier before lifecycle classification. Screening has not started; the qualifying count remains unknown, not zero of 127. One direct same-artifact episode that publishes all eight receipts would still improve the case evidence independently of that crosswalk.
Reader assistance can guide comprehension, not only answer questions
A randomized experiment with 117 higher-education participants compared three ways to support interpretation of a bar chart, communication network, and ward map: a conventional data story, a passive GenAI agent that answered questions, and a proactive agent that asked educator-authored scaffolding questions and gave feedback. All three groups improved from their unsupported baseline. After the assistance was removed, the proactive group’s median score was 6 of 6, compared with 5 for both the data-story and passive groups; the between-group effect was statistically significant and medium in the authors’ analysis. Completion time did not differ among interventions.
This is evidence for a technique, not a general conversational-chart mandate. The task covered knowledge and comprehension in one educational setting, used online participants with limited domain context, and did not measure newsroom, BI, decision, accessibility, or long-term field outcomes. Its useful lesson is specific: an assistant that structures attention and asks the reader to reason can produce a different outcome from one that merely supplies an answer.
AI-generated misleading charts can measurably damage comprehension
A 2026 controlled experiment tested 48 readers on chart questions in two phases. Both groups first read correct charts. In the second phase, the control group continued to see correct charts while the experimental group saw misleading bar and line charts produced through an automated AI attack framework. Accuracy was 88.3% in the control condition and 71.9% in the misleading-chart condition. After adjustment for baseline chart-reading ability, education, chart type, and misleading technique, the odds of a correct answer were 0.266 in the misleading-chart condition.
This directly measures one kind of reader harm: a data-consistent chart whose design induces wrong answers. It does not estimate how common such charts are, whether ordinary users or systems create them accidentally, or how they affect trust and consequential decisions in news, BI, health, or social media. The study used horizontal and vertical bar charts and line charts in a controlled question-answering task. Its useful warning is narrower and stronger than a general claim about “AI misinformation”: visual design can remain faithful to the underlying table while materially reducing reader accuracy.
Accessibility is a delivered interaction, not a checkbox
Two studies outside AI chart generation provide an important adjacent baseline for evaluating delivered artifacts. In a controlled smartphone study, 26 low-vision participants completed bar- and line-chart tasks under five conditions. The full interactive treatment—which combined space compaction, personalization, and selective viewing—had 100% task completion, compared with 61.5% using the baseline screen magnifier. On a simple bar comparison, mean time fell from 531 seconds with the magnifier to 235 seconds with the full treatment. The study assured accurate chart extraction and used experimenter-provided phones, so it does not establish performance on arbitrary AI-generated charts or readers’ own devices.
A 2026 study then compared screen-reader text, audio-tactile exploration, and a refreshable tactile display with 10 blind adults completing 360 task episodes. Device-level accuracy did not differ significantly, but completion time and workload did. Chart type mattered even more: accuracy was 88.9% for pie charts, 81.1% for bars, 63.3% for lines, and 44.4% for scatterplots across the studied systems. No single modality was best for every task.
Neither study tests AI assistance. Together they replace an empty accessibility row with measured acceptance criteria: test the actual device, representation, interaction, chart type, task, time, error, and workload. A generated text description, responsive layout, or nominally accessible export is not itself a reader outcome.
Readers judge the data-generating claim too
In a 117-entry public discussion of a chart about AI-generated web content, readers quickly challenged the reliability of the AI-content detector behind the data. The reaction is a case, not a representative study, but it captures a key reader behavior: people do not only inspect axes and colors. They ask whether the measurement method can support the headline.
AI-assisted visualization therefore inherits the full author-reader contract: source quality, transformation choices, uncertainty, visual encoding, intended claim, and publication context. “The chart is accurate” is too small a promise.
Different contexts require different forms of assistance
The same assistant behavior can help in one environment and damage another.
| Context | Primary human purpose | Useful assistance | Particularly dangerous shortcut |
|---|---|---|---|
| Journalism and public explanation | lead readers through evidence toward a bounded account | exploration before publication, annotation variants, accessible implementation, source and data checks | letting fluent generation substitute for reporting, authorial judgment, or audience testing |
| Operational dashboard | maintain awareness and support timely response | governed measures, anomaly explanation, stable layout edits, reader questions tied to source visuals | changing metric definitions or visual hierarchy without operational ownership and regression checks |
| Analytical BI | compare evidence and make decisions | semantic-layer grounding, query generation, visible intermediate data, reusable verification | confident answers over weak metadata, ambiguous measures, or hidden filters |
| Open-ended exploration | form and test hypotheses | many cheap views, branching, undo, direct manipulation, alternative transformations | prematurely turning a plausible pattern into a narrative conclusion |
| Scientific visualization | inspect domain-specific structures and support reproducible claims | restricted representations, executable transformations, linked views, domain checks | general visual plausibility standing in for specialized scientific correctness |
| Education and explorable explanation | help a reader build intuition | adjustable assumptions, guided interaction, multiple representations, feedback | generating interaction without measuring what learners actually understand |
| Presentation and one-off communication | communicate a deliberate argument in a constrained setting | rapid drafts, formatting, accessibility, export, speaker-supporting variants | producing generic polish without an intellectual trace of why this chart serves this audience |
A useful common spine can preserve the question, data, decisions, artifact, provenance, and acceptance evidence. It should then delegate authoring and evaluation to the context. There is no reason to force a newsroom narrative, an operational alert, and an exploratory notebook through one model of autonomy or one acceptance score.
The same lifecycle exposes different evidence holes
The common spine becomes useful when it makes unlike contexts comparable without imposing one acceptance threshold. The table below applies the same sequence—question, data authority, generation, checking, correction, delivery, reader use, and sustainment—to four environments where the evidence is strongest or the consequences are clearest.
| Environment | What evidence exists now | Where the lifecycle still disappears | Acceptance evidence that belongs to this context |
|---|---|---|---|
| Journalism and public explanation | Two current build diaries expose source work, task decomposition, data and geographic defects, correction, mobile checks, publishing, and proposed updates. | Independent verification, editorial review, reader comprehension, accessibility, later refreshes, and civic use were not observed. | Source-to-claim trace, editorial acceptance, delivered desktop/mobile/access states, audience comprehension, correction policy, and update ownership. |
| Governed BI and executive use | Product contracts expose semantic models, queries, permissions, and review surfaces. Public testimony and analytics essays explain why narrow metrics and visible assumptions matter. | Organization-owned definitions, routine task success, review queues, refusal, one-number incidents, trust recovery, and total cost are not independently measured together. | Approved metric contract, exact query and filters, permission result, human owner, refusal behavior, incident escalation, and decision outcome. |
| Operational software dashboards | Three same-artifact cases now cross public delivery into a later event. KubeStellar is a project-authored failure-to-recovery chain; OpenClaw adds a non-owner operator report plus an open repair proposal; Prism adds successive independent delivery failures, a corrected release, and a maintainer runtime check. Prism and OpenClaw also preserve accepted second-person AI-assisted change; Prism has a repeat contributor. | No case supplies independent acceptance and comparison, whole human-and-model cost, named maintenance-authority handoff, affected-user denominator, accessible-reader use, decision outcome, or calibrated trust. OpenClaw lacks an accepted repair and recheck; Prism lacks the affected operator’s final recheck. | Accepted change, repeat contribution, named authority, live route, configuration, incident and deploy chronology, affected users, unknown-versus-zero semantics, repair, corrected release, maintainer recheck, independent recheck, changed regression gate, total cost, and downstream outcome. |
| Learning and explorable explanation | A 117-person experiment measured immediate comprehension from proactive scaffolding; the classroom case above shows direct instruction rescuing a failed AI-assisted construction task. | Delayed construction, unfamiliar transfer, learning after ordinary assistant use, and a complete learner-to-reader project remain sparse. | Immediate and delayed unassisted performance, diagnosis, unfamiliar transfer, confidence calibration, and comprehension by the learner’s eventual audience. |
| Accessible and small-screen reading | Controlled non-AI studies measure low-vision smartphone and blind nonvisual task outcomes. An 11-participant MAIDR study covers the pre-AI implementation. A later abstract reports eight blind and low-vision participants using a legacy AI-study surface, and public code bounds that surface to the first containing tag, v2.10.0. Graphy adds the same three blind co-designers returning for 12 sessions across four workshops and eight months while the working interaction changed. |
MAIDR’s exact participant-tested build and model remain unknown. Graphy is a design study, not formal usability or performance evaluation; its workshops are not bound to immutable code or model snapshots, and its public repository has no tag, release, or representative recheck. | Keep co-design, formal evaluation, study build, first containing release, later changed release, and representative recheck as separate receipts; compare AI on/off with representative users on their own assistive paths. |
A lifecycle acceptance ledger
The direct answer from the current evidence is deliberately incomplete: no captured episode follows one AI-assisted visualization across all of accepted production, later maintenance, total human-plus-model cost, accessible use on readers’ own devices, a consequential decision, and calibrated trust. Three cases now cross public delivery and a real later event: one project-authored record reaches recovery, one non-owner operator report stops at an open failure, and one multi-attempt independent failure chain reaches a corrected release and maintainer runtime check but not an independent final recheck. The other cells remain missing.
A project history is not a same-artifact afterlife
A new horizontal audit tested four promising joins and admitted zero new episodes. The strongest near-miss is useful precisely because its parts must stay separate.
The CHI 2024 MAIDR study reports an IRB-approved evaluation with 11 blind participants across bar plots, heat maps, box plots, and scatter plots. Participants used their own screen readers and refreshable braille displays; the study retained task logs, interviews, and System Usability Scale means. That is substantive representative accessibility evidence for the studied pre-AI implementation.
The archived MAIDR repository
contains a dedicated later AI-study surface. Its final public introduction has
one active horizontal box-plot task at commit
125914d,
and v2.10.0 is the
first subsequent public tag found to contain that state. The
paper abstract reports eight blind
and low-vision participants and headline findings about
modal preference, LLM customization, and multimodal representation. Neither
the abstract nor the repository identifies the participant-tested build or
model snapshot. It is a version floor, not the tested version.
The legacy changelog
records later model swaps, verification-flow fixes, and releases through
v2.30.1, followed by deprecation. No selected event carries a representative-
user recheck. The current MAIDR project is a complete
TypeScript rewrite on a separate Git root, so it is a new validation target;
project lineage is real, but validation inheritance is not.
Repeated co-design is a lifecycle lane, not a release afterlife
The 2026 Graphy co-design study adds a different kind of longitudinal evidence. The same three blind co-designers completed 12 individual sessions across four workshops over eight months. The team reviewed transcripts after each round, selected changes, and implemented them before the next workshop. The interaction evolved from a continuous overview and free touch toward a layered presentation and a more explicit select → confirm → ask → verify sequence. In the final own-data workshop, all three used the complete interaction range without assistance.
That is stronger than a one-session preference walkthrough. It shows the same people returning after gaps, remembering and revising an interaction, and shaping a working artifact over time. It does not show formal efficacy. The authors explicitly describe the work as design knowledge rather than a formal usability or performance evaluation; agent accuracy and robustness were not quantified, filtering was still researcher-mediated, and controlled and in-situ evaluation remain future work.
The linked Graphy repository adds mechanism and change receipts, not the missing version join. Its seven public default-branch commits include prompt, query, state, orchestration, chart-loader, model-default, and Unity-crash changes. It has no public tags or releases, binds no workshop to an immutable commit or model configuration, and contains no representative recheck of current main. Current code cannot be projected backward into participant exposure.
Count Graphy as a separate longitudinal co-design lane. Do not count it as a complete release-afterlife episode. The shortest defensible answer is: three returning co-designers over eight months; zero formal release recheck. Reopen the lifecycle admission when a workshop is bound to immutable code and model state, or when a versioned release receives formal or in-situ representative recheck after a material change.
The full release-validation chain is still empty
A stricter audit asked whether any leading case joins six receipts on the same artifact lineage: repeated representative use, immutable tested build, exact exposure model, versioned release, later material change or event, and a representative or actor-separated post-change recheck. The answer is zero of seven in the named paper, repository, release, issue, and follow-up surfaces reviewed through 15 August 2026.
| Candidate | Strongest held receipt | First blocking join |
|---|---|---|
| Graphy | The same three blind co-designers return across 12 sessions and eight months while the interaction changes. | No workshop commit or exact model, no tagged release, no released-system recheck. |
| MAIDR-AI legacy lineage | Eight BLV participants are reported at abstract level; code adds a study surface, version floor, and later maintenance. | No participant-tested build or model and no representative recheck after maintenance. |
| Prism | One operator tests several public repair attempts; v1.8.11 is release-pinned and maintainer-checked. | No exact model exposure or affected-operator recheck of the corrected release. |
| OpenClaw Agent Dashboard | A tagged release precedes a non-owner operator failure and a separate non-owner repair proposal. | No merged repair, corrected release, reporter recheck, or representative-user denominator. |
| KubeStellar Console | One commit-pinned dashboard crosses live regression, repair, restored production, and prevention. | No exact model, repeated representative use, or actor-separated recovery check. |
| PM4Py-UCM | Development is frozen at v0.7.4 across 18 agent sessions, three model generations, 151 commits, and 20 releases. |
No representative use, elapsed post-delivery event, or separate user recheck. |
| DV-World | A released benchmark tests 50 native repairs and 80 changed-data or changed-requirement visualizations. | No delivered product release, elapsed event, or returning representative user. |
Every receipt type appears somewhere in the corpus. They do not add up across rows. Combining Graphy’s returning co-designers, MAIDR’s accessibility study, PM4Py-UCM’s frozen build ledger, Prism’s repair releases, and KubeStellar’s production incident would create a lifecycle no person or artifact actually experienced.
Use the six receipts as a release card. For a practitioner, the question is “did intended users return to the exact changed release I can use?” For an analytics leader, it is “which release cleared the representative recheck?” For a researcher or builder, preserve the code, model, configuration, release, change, actor, and outcome coordinates. For accessibility work, ask people on their own access paths to return after the AI-enabled release changes. For an editor or sponsor, the short answer is: 0/7 complete chains; fund the missing join, not another disconnected stage.
This is a bounded null, not proof that no case exists anywhere. Reopen it when one primary record binds participant exposure to exact code and model state, then follows a released version through a named change and a representative or actor-separated return check.
A second candidate exposes a different failure. A 2026 public-health copilot paper describes a 16-professional within-subject usability method with interpretability, trust, and usability ratings, but its results text says the separate human-subject experiment was removed and reports no human outcome. A participant method is not a result; a trust question is not calibrated trust.
Use seven keys before counting a new lifecycle episode: exact artifact and version, AI assistance on that artifact, accepted or released state, a later event, actor separation or independent recheck, sufficient source custody, and a genuinely new later-life outcome. A project name, elapsed study duration, feature documentation, or inaccessible paper record cannot fill a missing key. The bounded search does not show that no qualifying case exists anywhere; it shows that these four candidates do not supply one.
A new adjacent case makes the lifecycle shape more concrete. A 2021–2025 HealthTech visualization program followed 21 projects using 84 recorded calls, notes and decision logs, backlogs and issue trackers, one governance-dashboard case, and a 16-startup survey. The case reached deployment with patients. Dashboards remained useful when evidence and responsibility were attached to real decision moments; the authors also warn that the infrastructure requires continual feedback and maintenance. This is strong evidence about longitudinal governance visualization. It is not a study of generative chart authoring, and it does not measure the missing cost, accessibility, reader-decision, trust, or comparative-maintenance outcomes.
A prepared change test is the rung before maintenance
The 2026 DV-World benchmark adds evidence between first creation and field afterlife. Its 260 tasks include 50 native Excel chart repairs and 80 cases that adapt reference visualizations to new data, schema mutations, or changed requirements across Python, Apache ECharts, Vega-Lite, D3.js, and Plotly.js. The best reported agent reached 48% success on repair and a 51.44% overall evolution score. Human results were 88% on repair and 82.11–88.46 across the five evolution frameworks.
This is a meaningful change test. The released Excel evaluator compares broken, candidate, and gold workbooks through native Excel objects and hard-gates the chart components that require repair. The evolution evaluator combines gold-image rubric scoring with a value-frequency comparison over the candidate and gold tables. The paper’s error analysis includes destructive regressions during targeted repair.
It is not an observed maintenance record. The benchmark prepares the starting state and gold answer, then scores one candidate. Its “longitudinal” example is a two-stage port-and-refine trajectory, not a return to a delivered artifact after elapsed time. There is no real deployment, incident, refresh, owner change, second maintainer, reader outcome, or whole human-plus-model cost.
Use that boundary differently at each decision point:
- Practitioner: preserve a change ticket, before/after artifact, data comparison, render, and destructive-regression check before delivery.
- BI or analytics leader: add a changed measure, broken binding, refresh, and owner handoff to the pilot; compare time to accepted repair with the current direct workflow.
- Researcher: extend the benchmark with an independently introduced later event, retained trajectories and costs, and real maintainers or readers.
- Journalist or editor: revise a number, source, label, breakpoint, and accessible alternative after acceptance, then require another person to republish it.
- Product or engineering team: report accidental changes, turns, dollars, validations, and human intervention alongside the final score.
- Educator or access specialist: test diagnosis and recovery, then test the delivered result with intended readers on their own devices.
For an executive reader: current agents can now be tested on visualization repair and change, but the best reported systems still miss about half the benchmark and the study does not follow the artifact into production maintenance or count the whole cost.
Three same-artifact afterlives expose different missing stages
The boundary moves again with the 2026 KubeStellar Console experience report and its commit-pinned quality record. One maintainer reports building and maintaining the open-source Kubernetes dashboard with Claude Code and GitHub Copilot over four months. On April 2, AI-assisted PR #4239 passed the reported build and PR checks but introduced a runtime-only Vite substitution that blanked the live site. A user report exposed it. PR #4248 removed the unsafe substitution; PR #4253 removed a separate TypeScript build blocker so the fix could deploy. The project reports restoration in about 45 minutes and records a manually added post-build check for that failure class.
This is the first self-reported same-artifact production-afterlife case in the review. It preserves the difference between a code fix, a deployable build, restored production, and a future regression gate. It does not establish a comparative incident rate, lower total cost, safe autonomous maintenance, a successful handoff, accessible-reader use, or a better decision. The report is one solo-maintainer case with survivorship bias; its more autonomous Level 6 phase was about one week old, which the author says is insufficient for long-term conclusions.
A second open-source case improves the observer boundary without completing
the chain. The OpenClaw Agent Dashboard
README
says the cost, usage, rate-limit, session, and system-health dashboard was built
with Claude Code and documents local, service, and Docker installation. Its
v3.0.0 release
was published on March 5, 2026. On March 18, a non-owner operator reported
that the cost view showed $0 for all sessions
on a custom provider even though token data existed. The commit immediately
before the issue contains the provider/model lookup and zero fallback the
reporter identified. The owner acknowledged that self-hosted providers had not
been tested; the reporter followed up three days later, and the captured issue
remains open.
This is the second same-artifact afterlife and the first here with an independent operator-reported later failure. It reaches public distribution, third-party installed use, failure, and acknowledgement. It does not reach an accepted or merged repair, corrected release, or reporter recheck. A different non-owner’s open PR 31 proposes a provider-inference fallback matching the reported mechanism and states four local checks. No maintainer review or release follows. A patch is not an independent acceptance contract, and project-level “built with Claude Code” attribution does not establish that Claude generated the defective line. One configuration cannot establish a failure rate, invoice error, decision harm, or comparative maintenance cost.
A third case adds repeated repair attempts and makes verifier identity inescapable. The commit-pinned Prism README says the self-hosted family dashboard was built entirely with Claude Code, with the human maintainer directing requirements, design, and product decisions. That is project-level attribution, not evidence that the model caused any later defect.
In issue 81, a non-owner operator asked for a Home Assistant route because the existing Docker/nginx/certificate setup was too high a bar. The operator then reported an install failure, a failed repair build, and—after v1.8.7—install success but startup failure. They used Prism through ordinary Docker plus Pangolin instead. That route substitution is neither recovery of the Home Assistant route nor abandonment of Prism.
The later v1.8.11 release names missing Postgres socket-directory and non-idempotent schema-replay repairs. Its release-pinned entrypoint implements socket creation, first-boot schema gating, migrations, and failure log output. The maintainer reports installing, starting, and restarting the add-on on real Home Assistant OS. No comment from the affected operator or another independent installer follows. Prism therefore reaches a corrected release and maintainer runtime recheck, not independent recovery acceptance.
Keep five receipts separate before projecting one word such as “fixed”:
| Receipt | Minimum evidence | Prism issue 81 |
|---|---|---|
| Report | actor, version or route, environment, observed and expected behavior | present across successive independent add-on failures |
| Repair | named changed mechanism and commit or code path | present in release descriptions and the v1.8.11 entrypoint |
| Corrected release | public version containing the repair | present through v1.8.11 |
| Maintainer recheck | runtime, environment, exercised transitions, and result | present for install, start, and restart on real Home Assistant OS |
| Independent recheck | affected reporter or separate operator exercises the corrected release | missing after v1.8.11 |
An honest status is therefore released and maintainer-verified; independent recheck missing. Issue closure, a tag, code inspection, maintainer runtime verification, and independent acceptance remain different receipts.
Contribution needs its own state ladder. Prism’s non-owner PR 23 proposes a substantial weather widget and precipitation chart, explicitly credits Claude Sonnet 4.6, and preserves owner-found compose, test, and upgrade defects. The contributor repairs them; the maintainer makes additional integration changes; v1.7.0 credits the contributor; and the same person returns 17 days later with accepted PR 41. OpenClaw’s non-owner PR 15 explicitly credits Claude Code and enters its multi-provider pricing path after maintainer integration. KubeStellar’s held incident PRs remain creator-local.
| Handoff state | Minimum receipt | Current cases |
|---|---|---|
| Second-person change | non-owner changes the exact artifact and the owner accepts it | Prism PRs 23/41; OpenClaw PR 15 |
| Repeated contribution | the same non-owner returns after elapsed time with another accepted change | Prism, once after 17 days |
| Maintenance authority | a named second person can release or respond and later exercises that role | missing |
| Independent recovery acceptance | affected or separate operator exercises the corrected release in the affected route | missing for Prism issue 81 and OpenClaw issue 30 |
This is partial changeability evidence, not a transfer of operational ownership. Preserve the review defects, contributor repair, owner integration, released version, and later return; do not compress them into a contributor count or the word “handoff.”
Production recovery needs twelve actor-separated states
The three cases become more useful when their shared production edge is recomputed as one state model without combining rows. Every case crosses delivery into a later event; only KubeStellar and Prism reach accepted repair, corrected delivery, and maintainer recheck. No affected user or separate operator exercises either corrected state, and no accepted contributor is shown exercising release or incident authority.
| Production state | Held result | Decision boundary |
|---|---|---|
| AI-assisted provenance | 3/3 | project-level disclosure does not attribute a defect line |
| Pre-event release or delivery | 3/3 | name the exact route and version where available |
| Actor-separated exposure | 2/3 observed; KubeStellar partial | a project-retold user report is not a direct user receipt |
| Later production event | 3/3 | preserve the elapsed event outside the initial build |
| Diagnosis or reproduction | 2 observed; OpenClaw partial | plausible source mechanism is not a reproduced repair |
| Accepted repair | 2/3 | an open proposal is not an accepted fix |
| Restored or corrected delivery | 2/3 | merge and release remain separate from restored service |
| Maintainer recheck | 2/3 | name environment and exercised transition |
| Affected-user or independent-operator recheck | 0/3 | owner verification cannot fill this state |
| New prevention gate | 1/3 observed; Prism partial | name the check and later prove that it still runs |
| Accepted second-person change | 2/3 | contribution demonstrates changeability, not duty |
| Transferred maintenance authority | 0/3 | a named second person must hold and exercise release or incident authority |
The short production result is therefore 3 afterlives · 2 maintainer-verified restorations · 0 independent recoveries · 0 authority transfers · 0 complete twelve-state rows. OpenClaw issue 30 and PR 31 remain open. Prism issue 81 is closed after the maintainer’s v1.8.11 check, but the affected operator does not return. Ticket status changes no evidence state by itself.
A second 2026 single-case process-modeling experience report measures a different cost fragment: about 65 active hours across 18 agent sessions and ten weeks, 151 commits and 20 releases, with fixes outnumbering features 2.3 to 1. Visualization and layout work was 78% fixes. It makes refinement effort visible but does not follow a later field incident, handoff, reader, or whole monetary cost. Read the experience records together as study-design inputs, not as one synthetic episode.
Use this six-field admission card before calling a later record maintenance:
| Field | Minimum receipt | OpenClaw status |
|---|---|---|
| One artifact | persistent repository, route, report, or data-story identity | Yes: one named repository and dashboard |
| AI contribution | disclosure, transcript, or commit attribution tied to the artifact | Partial: project-level Claude Code disclosure, no per-change trace |
| Pre-event delivery | accepted version, release, deploy, or third-party installed use | Partial: tagged release plus one later third-party latest-clone installation |
| Later event | dated user report, incident, update, handoff, or retirement outside a benchmark | Yes: custom-provider cost-display failure |
| Same-artifact link | version, commit, route, file, issue, or owner record connecting the events | Partial: same repository and code path; reporter’s exact commit is not pinned |
| Disposition | repaired version and recheck, or explicit unresolved/retired state | Partial: owner acknowledgement and open repair proposal; no accepted merge, corrected release, or recheck |
The audience translations are different decisions over the same evidence:
- Practitioner: preserve accepted commit, contributor and owner changes, live route, configuration, incident, repair, corrected release, maintainer recheck, independent recheck, and the new gate beside the artifact; render a missing estimate as unknown rather than zero.
- BI or analytics leader: count contribution, named maintenance authority, monitoring, review, failed deploys, incident response, and ownership; require a real refresh and handoff before calling a pilot maintainable.
- Researcher: measure actor switches and reproduce the full chain prospectively with a direct-work comparator, total human and model cost, a second maintainer, and real readers; do not count acknowledgement as recovery.
- Journalist or editor: separate contributor change from editorial and production authority; treat a merged correction as incomplete until the public graphic is rebuilt, redeployed, checked on its delivery surfaces, and handed to another person.
- Product or engineering team: model contribution, authority, and recovery as separate states; link agent action to commit, checks, deployment, runtime symptom, repair, corrected release, maintainer recheck, reporter recheck, restoration, and changed prevention; distinguish unknown, zero, and not applicable, and keep human approval visible.
- Educator or accessibility specialist: teach collaborative verification across time and test the delivered result with disabled people on their own devices; automated accessibility checks are not reader outcomes.
For an executive reader: three AI-assisted dashboard codebases have observable afterlives and two reach maintainer-verified restoration. None has an affected-actor recovery receipt or transferred maintenance authority. This is evidence of production afterlife and partial changeability—not proof of a complete recovery, lower total cost, safe autonomy, or better reader decisions.
Use one ledger for the evidence, then read the decision layer that belongs to your role:
| Stage | Evidence to retain | Practitioner | Team or editor | Research and product |
|---|---|---|---|---|
| Frame | intended audience, decision, source authority, local definitions | Can I explain the claim and reject an incoherent request? | Who owns the metric, editorial judgment, and consequences? | Were the task, stakes, expertise, and baseline declared before use? |
| Build and correct | prompts, transformations, direct edits, failures, waits, rollback, rejected output | What did I inspect or repair rather than merely regenerate? | What review changed the artifact, and what remained disputed? | Measure time to first candidate separately from time to accepted work. |
| Accept and deliver | named approver, acceptance contract, exact published state, desktop/mobile/access checks | Did the artifact pass the real surface rather than only my preview? | Is approval distinct from publication, and publication distinct from reader use? | Recompute claims and capture the delivered state independently. |
| Use | defined reader tasks, comprehension, decisions, confidence, errors, feedback | What did intended readers understand and do? | Did the work support the decision without hiding uncertainty or excluding readers? | Measure decision quality and calibration, not satisfaction alone. |
| Sustain | refresh host, dependencies, second maintainer, incidents, correction policy, retirement | Can someone else update, repair, or retire it? | Who owns the next refresh and response when the artifact is wrong? | Follow at least one dependency change, handoff, scheduled update, or incident. |
| Account | human time, model or subscription cost, verification, failed generations, delivery, later maintenance | What was the whole cost after the impressive first draft? | Was the gain durable enough to justify adoption? | Compare against current direct work with the same scope and acceptance bar. |
Do not fill a missing cell with evidence from another project. Cases can teach the acceptance contract collectively; only a row observed from end to end can establish a complete episode.
What current products reveal about the direction of travel
Product documentation cannot establish quality, but it shows where vendors are placing control and context:
- General chat products now combine file upload, executed analysis, interactive tables, common chart types, limited direct customization, and download.
- Spreadsheet assistants preserve the installed grid while adding executed analysis and charts, but their refresh contracts differ materially: an inserted chart may be static or linked to copied rather than original data.
- Artifact builders turn prompts into editable, shareable applications rather than returning only prose or a static image.
- Guided editors and visible canvases bind the assistant to explicit settings or add new versions beside existing work, making natural language a shortcut into a state the user can inspect.
- BI and data-cloud assistants increasingly divide roles: authors maintain models, instructions, and permissions; analysts inspect queries and steps; consumers ask questions over curated data; administrators monitor usage and feedback.
- Open-source systems and skills package declarative schemas, library-specific syntax, tests, and production components so a stronger baseline model still operates inside a maintained execution contract.
- Research systems pair natural language with direct encodings, visible transformed data, history, branches, deterministic perceptual checks, and proactive reader scaffolding.
This is a meaningful convergence: natural language is becoming one control surface inside a structured environment, not the whole environment. As baseline models improve, generic “make a good chart” instructions are likely to lose relative value. Maintained business semantics, provenance, deterministic tests, edit locality, delivery evidence, and reader-specific acceptance are less likely to be absorbed by the model because they belong to the situation, not the pretrained baseline.
A practical way to evaluate a tool or workflow
Do not begin with “Which model made the prettiest chart?” Begin with a real job and measure the entire path.
- Name the environment and audience. State whether this is exploration, governed BI, operational monitoring, public explanation, science, or a reusable application. Name who must read or maintain the result.
- Choose representative work. Use self-owned data and include awkward semantics, missing values, multi-step transformations, a correction request, and the actual delivery surface. Keep a simple task as a control.
- Record the baseline. Measure current human time, error rate, correction path, and delivery effort. A tool cannot be said to save time when the counterfactual is unknown.
- Separate first draft from acceptance. Record time to first useful candidate and time to accepted artifact, including prompt writing, waiting, cleanup, verification, and publishing.
- Inspect consequential choices. Can the person see the fields, filters, aggregations, transformations, code or semantic query, and generated interaction state? Can they make a precise local edit without regeneration?
- Test the delivered artifact. Recompute values, exercise controls, check accessibility and responsive behavior, and verify that exported or embedded output preserves the intended state.
- Test readers separately. Ask defined readers to state the main claim, supporting evidence, uncertainty, and next action. Measure correctness, time, confidence, and harmful misreadings. Do not use a model grader as the reader sample.
What remains missing—and what the new evidence partly fills
The gap list is no longer uniformly empty. New controlled studies now quantify one correction setting, one form of AI-generated chart harm, and several mobile and nonvisual reader outcomes. A production-workflow interview study adds a high-stakes dissemination boundary. None closes the central distance between a controlled task and routine delivered work.
| Question | What can now be said | What remains open | Highest-value next study |
|---|---|---|---|
| Can AI critique support work over time on self-owned material? | Visualizationary followed 13 designers using self-selected data and tools over a three-to-five-day window and found average expert-rated improvement. | observed work was roughly 90–150 minutes per participant, with no baseline, production delivery, or later maintenance | multiweek within-person study on real commissions from intake through publication and update |
| What does correction cost, and when do people abandon? | in 108 forced-error task episodes, 7 were not completed and participants prematurely declared completion 31 times; the novice study measured failed repairs; Visualizationary recorded 68 feedback uses and selective rejection | the controlled tasks were engineered to fail and stopped at 15 minutes; there is still no field distribution of correction time, regeneration, rollback, abandonment, or downstream harm | instrument every prompt, edit, wait, undo, verification, and abandonment against a direct-work baseline |
| Do local semantics improve correctness? | verification research and current product contracts show why fields, measures, intermediate data, code, and verified queries matter; current BI cases report better trust with narrow curated metrics | no independent head-to-head field test of a general assistant versus the same model with organization-owned semantics | blinded task set with known local definitions, realistic ambiguity, and exact semantic-error scoring |
| Can the result be delivered and maintained? | 17 biomedical-visualization practitioners described using AI chiefly for auxiliary tasks and keeping substantial generated imagery out of final scientific work; two 2026 newsroom diaries expose publishing, mobile checks, architecture, proposed updates, and unfinished maintenance; Graphy follows three co-designers through 12 sessions and implemented changes across eight months; one direct second-maintainer episode exposes a creator-local refresh job; a four-year HealthTech program traces governance visualization through decision work and one patient deployment; DV-World adds 50 prepared Excel repairs and 80 new-data or evolving-requirement tasks; KubeStellar adds one project-authored same-codebase regression-to-recovery chain; OpenClaw adds a non-owner operator-reported cost-display failure plus an open repair proposal; Prism adds successive independent delivery failures, a corrected release, maintainer runtime verification, and an accepted repeat external contributor | Graphy is longitudinal co-design, not formal evaluation or a versioned release afterlife; the interviews measured judgment rather than delivery; the diaries, handoff, and KubeStellar case are self-reported; OpenClaw has no accepted repair or recheck; Prism has no independent final recheck; accepted outside contributions do not establish maintenance authority; the longitudinal program is adjacent governance visualization; DV-World is a gold-scored benchmark; no row joins whole cost, maintenance-authority handoff, accessible-reader use, consequential decisions, and calibrated trust | follow one case prospectively from prepared change and accepted delivery through an independently introduced event, repair, corrected release, maintainer recheck, independent recheck, changed gate, named second-person maintenance authority, intended-reader outcome, and whole cost |
| What is the total cost? | a 108-episode study measured task time and found no significant timing difference among conversational, stepwise, and phasewise interfaces; one 45-minute randomized exercise measured completion; the newsroom diaries add build durations, repeated correction, usage-limit waits, model-cost sensitivity, validation, and remaining work; PM4Py-UCM reconstructs about 65 active hours, 18 sessions, 151 commits, 20 releases, and fix-heavy visualization work across ten weeks; eleven specialist-ledger fragments now include one matched declared partial call/output-token budget | no common ledger joins human time, model inference or subscription cost, verification, failed generations, later delivery, maintenance, opportunity cost, and abandoned work against a current direct-work baseline; zero of eleven held fragments has both equivalent observed route-wide cost and a frozen accepted-artifact denominator | prospective cost diary retaining declared budget, every observed attempt and repair, cost per contract-accepted maintainable artifact, then one later event and restored delivery |
| Who benefits, by expertise and work context? | the human-skills review separates consumption, construction, critique, and connection from data, domain, tool, and delivery resources; the learning chapter adds a positive immediate post-removal comprehension result for proactive scaffolding, adjacent randomized evidence that assisted performance can outrun learning, and ordinary visualization-retention evidence showing faster procedural decay; bounded studies show experienced practitioners filtering critique and using constrained implementation effectively | “novice” remains inconsistently defined; expertise cells are small; delayed construction and far transfer remain mostly unmeasured; no captured longitudinal study causally estimates AI-driven visualization atrophy | preregistered capability-profiled field trial crossing answer-oriented and metacognitive assistance with delayed unassisted transfer in multiple job environments |
| Does assistance help a reader understand—or harm understanding? | the 117-person educational experiment found proactive scaffolding outperformed passive Q&A and a data story; a 48-person experiment found lower accuracy with AI-generated misleading charts than with correct charts | both are controlled question-answering tasks; no independent decision quality, public communication, operational action, calibrated trust, or durable field learning | independent reader trials in journalism, BI, education, and public services using delivered artifacts and consequential decisions |
| Does it work for disabled readers and on small screens? | a 26-person low-vision smartphone study measured large differences across delivery treatments; a 10-person blind-reader study found modality and chart-type trade-offs across 360 tasks; a 12-person AI-assisted learning study found strong preference and spatial-model benefit for tactile + text + chat, but no measured accuracy lift; Graphy adds three returning blind co-designers, 12 sessions, and implemented interaction changes across eight months; a 2025 position paper demonstrates information loss in model-generated description chains | the learning study was short and small; Graphy explicitly was not a formal usability or performance evaluation and has no workshop-version binding or released-system recheck; neither study evaluated delivered artifacts or isolated a chart-vision specialist; broader disability groups, readers’ own devices, screen-reader and keyboard behavior, responsive reflow, verification, and remediation remain open | bind every co-design session to code and model state, then run a co-designed AI-versus-non-AI delivery study on a versioned release using readers’ own devices, with task success, errors, time, workload, independent verification, recovery, and qualitative experience |
| What does the work feel like in context—and what are people trying to protect? | selected first-person accounts now distinguish mandate pressure, first-candidate momentum, repair interruption, reputational and scientific accountability, accessibility, craft, identity, and post-delivery disappointment; underheard-role evidence now includes ordinary spreadsheet allocation, provisional recipients, blind and low-vision learners, and one second maintainer | consequential recipients, disabled readers on their own devices, non-English workplaces, displaced workers, and multiple second maintainers across contexts remain too faint; public accounts cannot estimate prevalence | purposive multi-environment interviews and prospective diaries tied to actual artifacts, usage, corrections, decisions, handoffs, and abandonment |
| Which scaffolding still helps as models improve? | proactive questioning beat passive answering in one reader experiment; deterministic perceptual checks and structured interfaces supported bounded authoring; skill projects publish provider-run tests | few independent with-and-without ablations use the same current model, task, and harness, or repeat across model generations | recurring benchmark and field panel that removes one scaffold at a time and records conflicts as well as gains |
Semantic provenance, reader trust calibration, consequential decisions, and production regression remain especially thin. Harmful misinterpretation now has one controlled result, but no field estimate. The key experimental unit is not the generated chart. It is the complete creator-to-artifact-to-reader episode, with the job, environment, audience, and stakes named in advance.
Evidence guide: what the cited work actually is
| Evidence | Plain-language description | What it supports here | Important limit |
|---|---|---|---|
| 2024 State of the Data Visualization Industry | Online practitioner survey: 980 started, 763 completed, 825 answered the AI-use question; public respondent-level files permit role, task, and incumbent-tool cuts | observed entry points for AI in visualization work and directional differences among this sample | self-selected community survey with item-specific missingness; not representative adoption or product market share |
| 2025 State of Analytics Engineering | dbt survey of 459 practitioners and leaders on development, documentation, and natural-language data use | general versus specialized assistant use and the gap between analytics creation and conversational consumption | vendor-community sample, not visualization-specific, and based on self-report |
| DWP Microsoft 365 Copilot trial evaluation (2026 report) | Official mixed-methods workplace evaluation with 1,716 licensed-user survey responses, 2,535 comparison responses, and 19 quota-sampled in-depth interviews | routine work allocation, quiet non-use, Excel-specific limits, conditional trust, habit, time pressure, and stakeholder ownership | nonrandom licence allocation, no pre-trial baseline, self-reported outcomes, and visualization as a subset of a broad office-assistant evaluation |
| Vibe Visualizing (2026 preprint) | Think-aloud study of 20 visualization novices completing 60 ChatGPT chart tasks; 175 charts analyzed; initial prompts replayed on three current multimodal models | novice prompting, design failures, confidence, verification, repair, and changing model behavior | small controlled datasets; cross-model replay was not a live user study |
| How Good Is ChatGPT in Giving Advice on Your Visualization Design? (TOCHI 2025) | Rating study plus 12-practitioner comparison of ChatGPT and human visualization advice | brainstorming value, generic advice, context, expert preference | practitioner sessions used 2023-era GPT-3.5 |
| How Do Analysts Understand and Verify AI-Assisted Data Analyses? (CHI 2024) | Design-probe study of 22 professional analysts and 52 verification workflows | how people use code, data, explanations, and visual summaries to verify | prepared tasks at one company; not full chart authoring |
| Randomized public-health analysis exercise (2025) | 30 analyzed participants assigned to integrated ChatGPT or R/Stata plus ChatGPT for simulated epidemiological work | speed-quality separation, tool integration, serious errors | small, underpowered, simulated 45-minute exercise; no non-AI control |
| Data Formulator 2 (CHI 2025) | Mixed-initiative chart-authoring prototype studied with eight corporate participants | direct controls, visible transformations, history, inspection behavior | reproduction tasks, GPT-3.5, no baseline or long-term use |
| DashChat (2025 preprint) | Dashboard-prototyping system with formative interviews and a 28-person evaluation | pre-data prototypes, stakeholder alignment, structured edits | mockups rather than analytical correctness or production delivery |
| Visualizationary (2024 preprint) | Thirteen designers iterated self-selected visualizations with LLM and deterministic perceptual feedback over a three-to-five-day window; three experts rated change | longitudinal critique use, expertise differences, version tracking, and selective rejection of advice | small, short study using GPT-3.5; no critique baseline, production delivery, or later maintenance |
| GenAI agents and visual-analytics comprehension (2024 preprint) | Randomized 117-person comparison of a data story, passive Q&A, and proactive scaffolded dialogue with pre-, intervention-, and post-tests | direct reader-comprehension evidence and the difference between answering and guiding | controlled educational task, limited domain context and comprehension levels, no field deployment |
| Comparing LLM conversational and graphical interfaces for industrial decision tasks (2026 preprint) | Twenty participants used a dashboard and chatbot across four simulated tasks, followed by questionnaires and semi-structured interviews | conversational compression, dashboard overview and auditability, provisional reliance, and the separation between speed and confidence | exploratory proxy sample dominated by computer-science students, short low-stakes tasks, one implementation of each interface, and human-LLM-assisted qualitative coding |
| AI-supported end-user development for visualization (2025) | Eight interviews in one company followed by three profile-selected, 30-minute design-probe sessions | basic, intermediate, and advanced users’ different prompting, transparency, and control needs | exploratory single-company case; three sessions cannot establish usability or efficacy |
| How Do LLMs See Charts? (2026 preprint) | Comparison of 24 people and three multimodal models on 60 synthetic charts | why model and human chart reading are not equivalent | synthetic data and narrow designer-intent labels |
| Trustworthy by Design (2025 preprint) | 37 participants ranking charts and explaining perceived trust | clarity, source, familiarity, integrity, and aesthetics as trust cues | self-reported trust, not comprehension, correctness, or behavior |
| Vis Ex Machina (CHI 2021) | Preregistered experiment with 114 people choosing identical recommendations labeled human or algorithmic | contextual effects of provenance labels | predates generative charts; study charts were researcher-authored |
| Improving Steering and Verification in AI-Assisted Data Analysis (UIST 2024) | Eighteen experienced analysts completed 108 deliberately error-prone tasks across conversational, stepwise, and phasewise interfaces after a 15-person formative study | premature acceptance, non-completion, correction controls, task time, and the cost of structured intervention | forced-error 15-minute tasks using GPT-4 Turbo; no direct-work baseline, publication, or field abandonment |
| ChartAttack (2026 preprint) | Controlled 48-person comparison of correct and AI-generated misleading bar and line charts, with baseline chart-reading phase and adjusted analysis | direct reader-comprehension harm from selected visual misleaders | controlled chart QA and limited chart types; not prevalence, calibrated trust, or consequential decisions |
| Lexara (CHI 2026) | Interviews with 22 CVA developers, observation with 16 professional end users, and a two-week field-deployed diary in which six developers ran 38 experiments over 57 newly authored cases, ten models, and six system prompts | repeated diagnostic evaluation, model/prompt comparison, inspectable disagreement, and logged development-selection rationale | author-run study; no immutable participant-tested build, downstream accepted visualization, intended-reader decision, calibrated trust, whole cost, or later post-release recheck |
| Data visualization in AI-assisted decision-making, publisher supplement, and reconstructed DOI row (2025 systematic review) | PRISMA search across five databases through 1 July 2024; 127 reported studies, comprising 118 empirical papers and nine reviews; appendix domain totals sum to 127 | broad coverage, a published inclusion-inventory surface, a source-internal repair overlay for the ? row, and a complete 127-position stable-key register |
122 distinct A208 review keys; five exact-looking external DOI candidates lack review-controlled joins; one printed PMID resolves to another paper; lifecycle screening is unstarted and the qualifying count remains unknown |
| GenAI in biomedical visualization (TVCG 2026) | Ninety-minute workflow interviews with 17 biomedical-visualization designers and developers spanning research, production, and dissemination | where high-stakes practitioners use, restrict, or reject AI in real production workflows | purposive Euro-North American qualitative sample; no observed delivery or audience outcomes |
| Low-vision charts on smartphones (IEEE VIS 2024) | Fourteen formative interviews plus a controlled 26-person comparison of five ways to use bar and line charts on a smartphone | mobile completion, time, error, usability, and workload as delivered-reader outcomes | not an AI-generation study; accurate extraction was assured and participants used an experimenter-provided phone |
| Sound, Touch, or the Full Monty? (TACCESS 2026) | Ten blind adults completed 360 tasks using screen-reader text, audio-tactile interaction, and a refreshable tactile display across four chart types | modality, device, chart-type, time, accuracy, and workload trade-offs in nonvisual reading | not an AI-assistance study; small controlled sample with limited training |
| Touching or Chatting (2026 preprint) | Counterbalanced study of 12 blind or low-vision participants learning violin plots and clustered heatmaps with tactile + text + GPT-5.2 versus text + GPT-5.2; 263 substantive queries | direct AI-assisted accessibility experience, complementary spatial and conversational roles, preference, and the separation between mental model and accuracy | short small English-speaking US sample; no measured accuracy lift, long-term learning, or delivered-artifact evaluation |
| Conversational Tactile Data Interfaces (2026 author manuscript) and Graphy code | The same three blind co-designers completed 12 individual sessions across four workshops and eight months; the team implemented selected changes between rounds and all three used the complete interaction range unassisted in the final own-data workshop | longitudinal co-design, retention across gaps, evolving interaction mechanisms, and an open implementation lineage | not a formal usability or performance evaluation; no workshop-build or model binding, tagged release, representative recheck of current main, generalization, or whole cost |
| MAIDR: Making Statistical Visualizations Accessible with Multimodal Data Representation (CHI 2024), legacy AI-study surface, and current MAIDR project | Eleven blind participants tested the pre-AI system; a later abstract reports eight BLV participants using MAIDR-AI; public code establishes a legacy study surface and v2.10.0 version floor; current documentation declares a complete TypeScript rewrite |
representative accessible-use baselines, a source-custodied study locus, release floor, later maintenance, and an explicit rewrite boundary | no held source names the AI participant-tested build or model, detailed results remain abstract-only, later maintenance has no representative recheck, and the rewrite cannot inherit validation |
| Equity-aware generative AI copilot for digital public-health surveillance (2026) | Technical forecasting, fairness, anomaly, entity-agreement, and citation-coverage evaluation; the paper also describes a 16-professional usability method | a direct methods-results integrity example and an explicit field-deployment gap | the results text says the separate human-subject experiment was removed; no human trust, usability, real-decision, health-outcome, deployment, or maintenance result is reported |
| Second-maintainer dashboard discussion (2026 public testimony) | One direct handoff episode plus 91 exposed replies about an AI-built dashboard failing during its creator’s absence | hidden execution host, refresh lineage, semantic documentation, dependencies, ownership, and maintenance expectations | self-selected and unverified testimony; no prevalence, comparison, or causal estimate |
| Now You See Me (2026 preprint) | Embedded 2021–2025 governance-visualization program across 21 early-stage HealthTech projects, 84 recorded calls and operational traces, one dashboard case reaching patient deployment, and a 16-startup survey | decision-linked visualization, evidence reuse, role ownership, iteration, deployment context, and sustainment requirements | embedded UK/EU HealthTech program designed by the authors; not generative visualization authoring and no whole-cost, accessible-reader, calibrated-trust, or comparative-maintenance outcome |
| DV-World (ICML 2026) | Author-built 260-task visualization-agent benchmark with native spreadsheet creation/repair/dashboards, cross-framework evolution, and simulated intent clarification; four model trials per task and released evaluators | prepared repair and changed-data/schema/requirement stress tests, destructive regression, native chart binding, and a concrete rung beyond creation-only evaluation | benchmark-team evidence with no independent reproduction in custody; no accepted deployment, elapsed later event, real maintainer or reader, direct-work baseline, or whole-episode cost |
| AI Codebase Maturity Model / KubeStellar Console and quality record (2026) | One maintainer’s four-month experience report plus commit-pinned production QA and linked PR records for one open-source Kubernetes dashboard | same-artifact AI-assisted creation, hosted operation, user-detected production regression, two-step repair, restored deployment, and a new manual prevention check | self-reported single-maintainer case with survivorship bias; no independent comparator, whole cost, second maintainer, accessible-reader outcome, consequential decision, or calibrated trust; Level 6 observed for about one week |
| OpenClaw Agent Dashboard, v3.0.0 release, issue #30, accepted PR 15, and open PR 31 (2026) | Commit-pinned built-with-Claude project record, accepted non-owner Claude-assisted pricing contribution, tagged distribution, one non-owner operator’s later custom-provider cost-display report, matching code path, owner acknowledgement, and a different non-owner’s repair proposal | second-person accepted change plus the second same-artifact dashboard afterlife; explicit distinction among contribution, release, installed use, failure, proposal, acceptance, corrected release, and recheck | one configuration and no invoice ground truth, accepted post-incident repair, corrected release, recheck, whole cost, maintenance authority, accessibility, decision, trust, or causal attribution of the defective line to AI |
| Prism, PR 23, PR 41, issue #81, and v1.8.11 (2026) | Commit-pinned built-with-Claude project disclosure, accepted external Claude-assisted visual contribution through review and repair, a later accepted contribution, one non-owner operator’s successive Home Assistant failures, public repair releases, release-pinned repair code, and maintainer runtime recheck | second-person change and one repeat contributor plus the third same-artifact afterlife; explicit separation of contribution, maintenance authority, report, repair, corrected release, maintainer recheck, independent recheck, and route substitution | one operator; no independent final recheck, maintenance-authority transfer, exact action-level AI attribution, whole cost, representative accessibility, dashboard-informed decision, trust, or comparative rate |
| PM4Py-UCM experience report (2026 preprint) | Single expert-author case reconstructed from 18 agent sessions, 317 substantive human turns, 10,328 tool actions, 151 commits, 20 releases, and about 65 active hours over ten weeks | measured build and refinement effort for a process-modeling tool with dashboards; fix-heavy visualization and layout work; executable verification infrastructure | author is subject and analyst; no direct-work comparator, model-dollar ledger, later production incident, second maintainer, or reader outcome |
| Public practitioner and reader threads | Unrecruited discussions about current BI and spreadsheet work and chart reactions, including a 190-entry Excel discussion and a July 2026 dashboard-versus-chat discussion | concrete jobs, delights, friction, abandonment, changes in peer learning, trust repairs, and reader questions that occur in practice | self-selected testimony; establishes possibility of an experience, never prevalence or a population trend |
| Selected first-person creator logs and practitioner interviews | Named accounts spanning early experimentation, guided learning, personal dashboards, visualization coursework, analytics practice, data journalism, accessibility, and analog craft | how anticipation, delight, frustration, responsibility, refusal, and professional identity change across stages of work | editorially selected and unusually articulate accounts; self-reported episodes are not comparative tests, population estimates, or independent artifact audits |
| Two 2026 Reuters Institute build diaries | First-person accounts of rebuilding two data-journalism projects and constructing one national health dashboard with current coding agents | production sequence, time compression, visible defects, correction, architecture, mobile checks, validation, ownership, and unfinished work | three projects built and checked by their authors; no equal-budget baseline, independent audit, reader outcome, or maintenance follow-through |
| Playing Telephone with Generative Models (2025 position paper) | One source visualization passed through four GPT-4o and Claude Sonnet 4 description-and-image chains | mechanisms of information loss, fabrication, misplaced emphasis, poor design comprehension, under-description, overconfidence, verification disability, and compelled reliance | authors explicitly frame the work as a provocation rather than a controlled experiment; no disabled-reader study or error-rate estimate |
| Doom or Deliciousness (CGF 2023) | Semi-structured interviews with 21 experts in visualization or HCI, art or art history, and machine learning | early field expectations about help and harm across datafication, transformation, visualization, and interaction | convenience sample, mostly limited direct generative-tool experience, image-generation emphasis, and speculative framing |
| Dated field essays, interviews, and teaching or design cases | Role-diverse assessments from visualization researchers and educators, analytics and library practitioners, a blind journalist, a creative studio, and data journalists | purposes, workflow mechanisms, productive disagreements, lived constraints, and questions the controlled literature should test | perspective and self-report are not prevalence, comparative efficacy, or reader-outcome evidence; positionality and publication date remain attached |
| Current product and project documentation | Provider or maintainer descriptions for general assistants, spreadsheets, notebooks, editors, governed BI, data clouds, open-source systems, and skills | the feature, data-access, target-user, control, and delivery contracts available in August 2026 | self-reported feature scope and positioning, not independent adoption or comparative performance |
Scope and limits
This review is a moment-in-time synthesis, not a market-share study or product ranking. Twenty-four human studies and structured workplace evaluations and one 127-study systematic review were read in full, along with two current single-case AI-assisted development reports, one independently reported open- source dashboard failure chain, practitioner surveys, current provider and open-source documentation, public practitioner and reader discussions, first- person production cases, and dated essays and interviews from several professional positions. Study populations, tasks, models, and evidence vintages remain visible because results are not directly interchangeable.
Public threads were used to discover and illustrate experiences that controlled research often omits. They were not sampled to estimate prevalence, and quoted claims were not independently reproduced. Vendor documents establish what a product says it supports, not that the feature is accurate or useful. Several important human-computer-interaction studies use 2021–2024 systems; their raw model comparisons are dated, while their observations about context, verification, and human control remain relevant hypotheses for current tools.
The strongest counter-reading is that better 2026 models may erase many errors reported in earlier work. The fresh novice study partly supports that: replayed prompts produced fewer design flaws with newer systems. It also found new failure modes from richer outputs, including render failures, misleading interpretations, latency, unclear interaction state, and greater verification burden. Capability is moving; the location of the human work is moving with it.
Update log
-
2026-08-16 — Retired repeated B91 retrieval without classifying the paper. Three bounded lawful passes exhausted the same named surfaces, so generic search is paused with four exact reopen conditions. The content ledger remains
6 / 20 / 1 / 0; one access-blocked row remains unassessed, zero generic B91 targets remain active, and effort moves to a stronger lifecycle-bearing target unless exact new custody appears. -
2026-08-16 — Added the InfoViP actor and closed-opportunity afterlife. Three current research mentors now have bounded role coordinates; the closed listing, nonemployee fellowship contract, and adjacent QA project still establish zero named operating or approval authorities and no shipped InfoViP–Elsa integration.
-
2026-08-16 — Repaired the InfoViP episode chain instead of accepting search recency. The future-tense FDA page is content-current to 2024, not 2026. A current FDA biography names a project lead without establishing operations or maintenance authority, and the Elsa 4.0/HALO launch does not name InfoViP. The maintained chain remains 15 named cases and zero complete episodes.
-
2026-08-16 — The final InfoViP report adds component approval and production integration, plus explicit QA limits. It reports AWS/AERS integration, more than 30 million historical and about 8,000 daily submissions, fully implemented case-series use, and administrator monitoring. It also says the QA plan is absent, audits are planned, routine roles are incomplete, and signal-discovery effects are uninvestigated. A 2026 Elsa/API/UI opportunity is a prospective change trigger; 15 named cases still yield zero complete episodes.
-
2026-08-16 — InfoViP added counted real-work use, operating scale, and a bounded award envelope. A six-month internal evaluation records 20 unique reviewers and 58 submissions; a 2025 primary abstract reports operation over 29 million historic plus daily incoming reports; three FDA awards total $2.40 million obligated and $2.86 million estimated. Current routine use, whole-platform acceptance/ATO, outcome, affected-reviewer return, and whole cost remain absent; the audit stays at 15 named and zero complete episodes.
-
2026-08-16 — InfoViP crossed into an operational pipeline, not a complete lifecycle. Joined seven-evaluator co-design to an agency-reported pilot, 28-million-report plus incoming-report processing, a later performance and human-quality-control report, and current development/no-ATO counterevidence. The combined audit is now 15 named cases and zero complete episodes.
-
2026-08-16 — One publisher abstract closed one content gap, not a lifecycle. B92 now contributes a multicriteria, Monte Carlo, rank-stability decision-support method at the exact publisher-abstract layer. B91 remains content-unassessed; the current ledger is six full texts, 20 abstract or official-content rows, one unassessed row, and zero complete AI lifecycles.
-
2026-08-16 — Live source lineage tested rather than assumed. The current DiscoverWater page does not match the sole published page state. Two of three sampled data assets match canonically, but no exact build, audience-use, outcome, accessibility, cost, or AI receipt was recovered.
-
2026-08-16 — One public-delivery afterlife kept below use and impact. B110 gained an official abstract, pinned framework and application source, and a same-named DiscoverWater interface reachable at KU. The residual is now B92 and B91; no live-build join, continuous or ordinary-use denominator, audience outcome, recheck, whole cost, or AI role was admitted.
-
2026-08-16 — The residual primary-content gap reduced to three. All 14 remaining DOI rows were investigated; eleven gained substantive content, including one full chapter, while three remain metadata-only. One expert evaluation-to-redesign mechanism and two practice-context near misses were added without promoting any row into a GenAI role, accepted release, routine-use denominator, consequential outcome, or later recheck.
-
2026-08-16 — Seven human evaluations kept below one deployed lifecycle. Eight new primary abstract or official-project surfaces add seven explicit participant/evaluator studies. InfoViP is the strongest delivery near-miss, but its NLP and unsupervised-learning production installation is prospective; 14 DOI rows remain content-unassessed and zero complete lifecycles are added.
-
2026-08-16 — Field use kept below AI lifecycle. Repaired three DOIs, two publication years, and one foreign PMID across eight stable keys; added the cancer diary’s increasing-use and 11-clinician usability receipts while retaining its missing AI role, accepted build, outcome, later event, and affected-clinician recheck.
-
2026-08-16 — Human evidence separated into three pre-deployment lanes. Five exact full texts entered custody below A208’s 27 no-abstract DOI rows. Four empirical papers add controlled audience measurement, public-sector co-design/demo/bounded use, and enterprise demonstrator/production-intent receipts; none joins accepted field release to later affected-audience recheck. Twenty-two DOI rows and eight non-DOI keys remain unassessed.
-
2026-08-16 — Mature visualization afterlife kept outside the AI denominator. Exact-DOI discovery covers 114 stable keys and 87 available abstracts, adding no hidden GenAI signal. AIDSVu contributes ten years of public delivery, aggregate use, governance, and a later 2026 data release, but no AI role, immutable build, version-bound audience outcome, or affected- audience recheck. Eighty-six abstract nonmatches remain unexcluded.
-
2026-08-16 — One pinned output bundle kept below audience evidence. A bounded screen of A208’s 122 authoritative stable-key titles surfaced one explicit GPT study. Its one-commit supplement preserves 1,271 blobs across quizzes and homework, but exposes no immutable provider snapshot, named release, later commit, audience delivery, audience outcome, or recheck. The other 121 title nonmatches remain unscreened rather than negative.
-
2026-08-16 — Longitudinal development decision kept below audience consequence. Added Lexara’s two-week deployment with six CVA developers, 38 experiments, 57 newly authored cases, ten models, and six prompts. The combined audit now covers 14 named reader, production, and evaluation cases; zero joins all twelve receipts from immutable build through accepted delivery, consequential decision or calibrated trust, and post-release recheck.
-
2026-08-16 — Complete review register keeps five candidates outside the denominator. The six domains contribute 127 supplement positions and 122 distinct A208 review keys. Five exact-looking DOI identities lack review- controlled joins, one printed PMID points elsewhere, and screening remains unstarted rather than producing a false zero.
-
2026-08-16 — Education crosswalk stops at 12 of 13 authoritative keys. Eleven printed labels plus B59 resolve to A208 bibliography/DOI keys.
M. Lu (2020)is supplement-only relative to A208; a plausible Crossref paper is retained as a candidate, not silently admitted. The public account now separates source label, review key, candidate, and admission authority. -
2026-08-16 — Missing Education name reconstructed without rewriting the source. A208’s post-selection Education paragraph leaves B59, Hernández-Calderón et al. (2023), as the unique example absent from the twelve named supplement entries; its bibliography and Crossref agree on DOI
10.1093/iwc/iwac043. The public account now carriespublished: ?beside the repair overlay while keeping the full stable-key crosswalk and lifecycle count unfinished. -
2026-08-16 — Review inventory found count-complete but identity- incomplete. The publisher supplement’s six domain totals sum to 127, but its “full list” names 126 citations and leaves one Education entry as a literal question mark. Repair now precedes any complete row-level screen; the reported total stays 127 and the lifecycle count stays unknown.
-
2026-08-16 — Review breadth separated from lifecycle denominator. Added a 127-study decision-visualization review as a coverage map, retained its explicit non-GenAI and longitudinal limits, and kept the qualifying same-artifact lifecycle count unknown rather than reporting zero of 127.
-
2026-08-16 — Production recovery split into twelve states. Consolidated three dashboard afterlives into an actor-separated ledger: two reach accepted repair, corrected delivery, and maintainer recheck; zero reaches affected- actor recheck, authority transfer, or all twelve states.
-
2026-08-15 — Reader-outcome delivery boundary. Added an eight-receipt lineage test across ten primary rows. ChartAttack supplies the sole direct controlled AI-created-chart reader effect; zero row reaches accepted delivery, representative intended-context outcome, and later recheck.
-
2026-08-15 — Product controls separated from governed deployment. A seven-row governance audit now keeps three provider control contracts, two named-feature organizational-use stories, and two adjacent longitudinal governance processes distinct. Zero row carries the nine-receipt same-deployment chain; the gap is bounded to the held public surfaces.
-
2026-08-15 — Three controlled bridges, zero accepted-delivery bridges. Audited ten primary cases across technical evidence, human use or correction, acceptance, delivery, reader outcome, maintenance, and whole cost. HAIChart, interactive task decomposition, and ChartAttack form three bounded partial bridges; none reaches exact-version accepted delivery or a complete episode.
- 2026-08-15 — A fixed inference budget is not accepted-work cost. Added Selective TTS as the strongest held matched-budget near miss. Its declared calls and separate output-token budgets are close, but released accounting omits repair and failed-worker cost and its “final reports” are candidates. Practitioners now get three receipts—declared, observed, accepted—and the complete result is 0/11.
- 2026-08-15 — Six receipts, seven cases, zero full chains. Audited Graphy, MAIDR-AI, Prism, OpenClaw, KubeStellar, PM4Py-UCM, and DV-World against one exact release-validation join. Every receipt appears somewhere; no row joins all six, and the result remains scoped to named surfaces rather than the world.
- 2026-08-15 — Repeated co-design separated from release afterlife. Added Graphy’s three returning blind co-designers, 12 sessions, four workshops, eight-month span, implemented interaction changes, seven public commits, and explicit formal-evaluation limit. No workshop-version binding, release, later representative recheck, or complete lifecycle episode was added.
- 2026-08-15 — Accepted cost separated from attempt averages. Added the two-receipt rule for practitioner productivity: keep every attempt and failure in route cost, count only contract-passing artifacts, and report the then-current 0/9 held-comparison result without inferring a winner.
- 2026-08-15 — Version floor is not the tested version. Pinned MAIDR’s
later AI study to a dedicated legacy study surface and
v2.10.0as the first containing tag. The abstract reports eight blind and low-vision participants; the tested build, model snapshot, detailed results, and any representative recheck after nine selected maintenance and deprecation events remain missing. - 2026-08-15 — A project history is not a same-artifact afterlife. Audited four promising horizontal joins and admitted zero new episodes. Added representative pre-AI MAIDR accessibility evidence, retained the current AI-enabled rewrite as a separate artifact state, quarantined a described-but- removed public-health human study, and exposed the seven-key join for each intended audience.
- 2026-08-15 — Accepted contribution is not maintenance handoff. Added Prism’s external visual contribution through review, repair, release, and a later accepted return; OpenClaw’s accepted external pricing work; and its separate open repair proposal. Keep contribution, authority, and recovery as different receipts.
- 2026-08-15 — Repaired release is not independent recovery. Added Prism as the third same-artifact dashboard afterlife and the first held multi-attempt delivery chain. It reaches v1.8.11 and a maintainer install/start/restart check; the affected operator’s final recheck and the rest of the complete lifecycle remain missing.
- 2026-08-15 — A second dashboard afterlife, observed from outside the project. Added OpenClaw Agent Dashboard’s release-to-failure chain: one non-owner operator reported a misleading zero cost display, commit-pinned code supports the mechanism, and the owner acknowledged an untested configuration. Added a six-field admission card and kept repair, corrected release, recheck, whole cost, handoff, accessibility, decisions, trust, and AI causality open.
- 2026-08-15 — One same-artifact production afterlife. Added the KubeStellar Console incident chain from AI-assisted change through live regression, two-step repair, restored deployment, and a new prevention check; added PM4Py-UCM’s measured refinement-cost fragment; translated the result for six audience routes; and kept whole cost, handoff, accessible reader use, consequential decisions, calibrated trust, comparison, and transfer explicitly open.
- 2026-08-15 — Maintenance-shaped benchmark evidence. Added DV-World’s 50 native Excel repair tasks and 80 visualization-evolution tasks, audited the released evaluators, translated prepared change tests for six audience routes, and preserved the boundary between one benchmark episode and a real artifact’s afterlife.
- 2026-08-15 — Lifecycle evidence and audience decision layer. Added a four-year HealthTech governance-visualization program, preserved why it is adjacent rather than generative-authoring evidence, and introduced a shared lifecycle acceptance ledger translated for practitioners, leaders, editors, researchers, product teams, and accessibility work.
- 2026-08-14 — Initial public edition. Mapped practitioner jobs and tool choices, synthesized creator and reader accounts, and separated possible output, accomplished work, delivery, and lived outcome.