Can it make the thing?
A chart renders. A dashboard runs. A prompt produces an analysis or interactive artifact.
Evidence: execution, task completion, feature contractResearch snapshot · August 2026
AI is already good at getting someone from a blank page to a plausible first thing. It is much less reliable at carrying the work through local meaning, correction, delivery, and audience understanding.
The reliable win is compression, not delegation. Faster creation matters; it does not transfer responsibility for the analytical claim or the reader’s understanding.
For creators, readers, and product teamsThis companion translates research results into the lived jobs, gains, costs, and trust questions surrounding current tools.
Not a product rankingProducts appear to explain interaction environments. Documentation establishes features, not comparative quality.
The central gap
Most demos answer the first question below. Practitioners live in the second. Readers determine the third.
A chart renders. A dashboard runs. A prompt produces an analysis or interactive artifact.
Evidence: execution, task completion, feature contractThe data and semantics are right. Decisions are inspectable. Precise edits, delivery, and reuse are possible.
Evidence: accepted artifact, total time, errors, repair, maintenanceThe creator can explain the process. The reader grasps the claim, limits, and appropriate next action.
Evidence: observed behavior, reader tasks, decisions, confidence calibrationHAIChart connects recommendation performance to controlled analyst use, interactive task decomposition connects system error to analyst correction, and ChartAttack connects generated attacks to reader harm. None freezes an exact accepted artifact and follows it through real delivery; none completes the reader, later-maintenance, and whole-cost episode. These are three partial paths—not one synthetic lifecycle.
Who is reaching for what
The best available adoption evidence shows practitioners adding AI beside spreadsheets, code, design, and BI tools. Creation and analysis are ahead of trusted conversational consumption.
Online, self-selected survey. The result describes respondents, not the population or product market share.
Use was up 13 percentage points from the 2023 survey. The 28 unsure responses are the narrow violet segment.
Multiple selections were allowed, so the bars do not sum to 100%.
A general assistant beside Excel, Tableau, Python, R, D3, Figma, or a publishing tool
Prepare, debug, learn, storyboard, draft labels, make a first view
Observed survey patternSpecific AI-product share was not measured.
General coding assistance, then project-aware notebook or development agents
Write SQL or Python, document models, debug pipelines, create analysis output
Observed survey patternBroader than visualization and vendor-adjacent.
AI inside Excel or Sheets; AI spreadsheets when code or live connections outgrow the grid
Ask about a range, create formulas and charts, preserve a familiar handoff
Installed-base inferenceExcel use is observed; AI-feature adoption is not.
AI inside an existing BI, notebook, or governed data platform
Author reports, reuse measures, inspect queries, answer follow-ups, govern access
Provider target + casesNo independent head-to-head field test.
Conversational BI over curated metrics
Retrieve a number, ask why it changed, get an ad hoc cut without navigating a report
Provider target + testimonyTrust depends on narrow, owned definitions.
General help for ideation or code, then a deliberate design and publishing surface
Explore, create variants, annotate, improve accessibility, implement a story
Small subgroup + inferenceNot a population estimate.
DVS survey and public data ↗ · Analytics-engineering survey ↗. Both are self-reported and nonrepresentative; denominators remain attached.
The jobs
“Make a chart” hides work before, during, and after visual encoding. The same assistant can be strong at one job and harmful at the next.
Decide what is worth doing and which data is allowed to answer it.
Turn source material into evidence without losing its meaning.
Choose what the audience will see and make every consequential choice inspectable.
Test the real surface, support the reader, and keep the work correct later.
Listening to the field
The same person can be delighted by a first draft, frustrated by the repair loop, afraid to put their name on the result, and still choose to use the tool tomorrow. Attitude and behavior do not move together.
People arrive with organizational demands, prior skill, career hopes, and fears about what assistance may remove.
Leadership wanted visible thought leadership. The custom chatbot was unreliable; for the immediate task, a pivot table was faster. The frustration was having to perform adoption.
Jisell Howe first worried that instant code would remove the journey from idea to customized chart. A concrete build changed the feeling: errors remained, but search and troubleshooting became less disruptive.
AI produced a first bar chart, but an instructor supplied the visualization knowledge needed to improve it. The gain was reaching the work—not yet independent mastery.
Delight tends to attach to access, flow, and continuity across chores, not only to a chart appearing.
A two-day revision felt surprisingly fluid. Four short lines were enough once the current files, requested changes, and source URLs were present. The creator’s lesson was that organized context mattered more than prompt polish.
Sef Kloninger called the work “plain fun.” The agent moved with him from a 500 GB dataset through tests, exploration, a dashboard, debugging, and a presentation; he still inspected raw data and requested checks.
The practitioner immediately added the caveat: prior proficiency made the compression possible. A novice in another account valued a smaller gain—producing cleaning and validation scripts that had been out of reach.
The experience depends on whether explaining, waiting, inspecting, and retrying costs less than direct manipulation.
Stale filters, bookmark identifiers, whitespace churn, and slow edit-preview cycles accumulated. One experienced practitioner said describing a small change could take longer than making it.
Christina Stathopoulos described useful brainstorming and exploration, then recalled a generated chart whose accompanying interpretation reversed the visible comparison. The feared failure was looking careless in front of a stakeholder.
A recognizable site appeared in about an hour. Geography, evidence notes, rate limits, mobile behavior, architecture, and fact-checking then consumed repeated rounds of repair.
The imagined stakeholder, patient, client, colleague, or excluded reader changes which errors matter and where refusal is rational.
An agent recreated roughly a week’s visible dashboard work from a finished screenshot and source tables. The consultant’s question was what clients had really paid for: construction, diagnosis, or the judgment embedded in the reference.
Participants welcomed boilerplate, inspiration, and translation while some rejected final AI imagery for anatomy, patient communication, or scientific work. Others protected rendering because it was also where control, flow, and creative joy lived.
Johny Cassidy describes exclusion from charts with missing or useless descriptions—and the relief of colleagues taking responsibility. Generated alt text is not accessible delivery without alternatives, feedback, and correction.
Use, maintenance, correction, learning, and second-person handoff determine whether the saved effort became value.
A manager praised an urgent dashboard and sent it to six colleagues. Three months later, the usage record showed one manager view. Creator pride, stakeholder approval, and reader use were three different outcomes.
AI documented SQL and calculations, reviewed junior work, and rehearsed stakeholder questions. Those tasks made the dashboard easier to explain and maintain without asking the model to own the final claim.
Emilia Ruzicka collected and drew personal data by hand. Slowness, imperfection, and direct contact with the data produced experimentation and care—the kind of learning a faster route can accidentally remove.
Listening past the original creator
A second pass pursued spreadsheet-native work, people who do not make AI their default, recipients deciding whether an answer is defensible, blind and low-vision learners, and the second person asked to maintain the result.
DWP staff allocated tasks according to expertise, time, trust, data sensitivity, and habit. Data-heavy Excel work, chart generation, and intricate formatting remained weak points. One person valued having the assistant but was often too busy to remember to use it.
Chat 15Dashboard 3Both 2
Chat 0Dashboard 18Both 2
A supply-chain analyst noticed that AI had displaced visits to a peer forum. Replies described both useful help and invented functions, wrong references, damaged formulas, and refusal. Several people valued the incidental learning produced by solving someone else’s problem in public.
Program managers and funders wanted overview, drill-down, definitions, neutral language, and review support. In a separate sales-dashboard demonstration, an observer rejected the idea that an executive seeking a quick answer should independently validate regenerated charts against raw data.
Eleven participants preferred tactile charts plus text and chat; one said the better mode depended on complexity; none preferred text and chat alone. Touch supplied spatial structure and chat supplied flexible clarification. Measured chart-understanding accuracy did not improve.
An AI-built dashboard failed while its creator was away. The inheriting maintainer replaced the creator-local scheduled job with a proper pipeline; replies surfaced missing metric semantics, hidden dependencies, documentation gaps, and service expectations nobody had owned.
What listening adds
Field perspectives
These sources expose purposes and production constraints that benchmarks omit. Most are first-person cases, interviews, or essays: they establish that a perspective or experience exists, not how common or effective it is.
Two current newsroom diaries describe work that once took weeks appearing in days or hours.
They also contain wrong totals, broken geography, missed notes, repeated correction, usage limits, mobile checks, and new validation work. Benn Stancil explains why a plausible chart cannot validate its calculation.
Semantic stakes, independence of the check, cost of error, and ownership of the complete pipeline.
Enrico Bertini maps opportunities across data acquisition, wrangling, analysis, tours, and creative exploration. Alberto Cairo treats easier code as capacity extension.
Question quality, interpretation, and final analytical choice become more important. Exploratory, explanatory, essayistic, and artistic work do not share one quality function.
Purpose, audience, consequence, capability profile, and whether the claim remains inspectable.
Removing syntax and formatting work can get a person to a candidate before the question goes cold.
A classroom exercise used failed prompting to reveal missing chart structure. An analog practice made slowness a source of attention, experimentation, and care.
Whether the effort is obsolete mechanics or practice in structure, skepticism, memory, and meaning.
Generated descriptions and alternatives may help newsrooms address access at a scale that manual workflows have not reached.
Elavsky and Xiong Bearfield show model chains losing data, source, uncertainty, design, and purpose. Blind journalist Johny Cassidy describes an organizational and multimodal problem, not an alt-text checkbox.
Whether the reader can test the account, choose another modality, report failure, and obtain remediation.
Richard Brath asks which charts remain useful when models extract insights. Elijah Meeks proposes audience, intent, interpretation, and conversation as framework objects.
Domestic Data Streamers argues that visualization remains a distinct language for comparison, human experience, and creative divergence.
Reader task, need for overview or shared evidence, qualitative meaning, novelty, and an auditable interaction trace.
Production cases
Both are self-reports by the people who built and checked their own projects. Keeping the gain beside the incurred work prevents “built in a week” from becoming a quality claim.
One map fell from roughly three weeks in 2012 to a two-day reconstruction.
A recognizable site appeared in about an hour; the full dashboard and checking workflow took one week.
An equal-budget comparison with the best current direct workflow or another practitioner.
Incorrect totals and incomplete bilingual text required source comparison and repeated correction.
Cities rendered in the sea or broken shapes; exact evidence and source notes were sometimes missed.
A field distribution of error, correction time, regression, premature acceptance, or abandonment.
About 20 ordered tasks, source checks, and open code made components easier to trace and revisit.
Extraction, data, presentation, insights, and fact-checking were separated so repairs could be replayed.
That decomposition or a model-built validator independently improves correctness across projects.
Working prototypes were published; attention, understanding, and civic use remained the stated bottleneck.
Mobile rendering was checked and stakeholders reacted; readers were not tested.
Audience understanding, accessibility, consequential use, or sustained adoption.
Record reconciliation and scheduled updates were proposed next steps.
Refactoring and longer-term maintenance remained after the one-week build.
Later refreshes, dependency changes, second-person handoff, incident response, or retirement.
Ottaviani build diary ↗ · Tordecilla build diary ↗. These cases support process claims, not general productivity or accuracy estimates.
The tools
This is a showcase of current approaches, not a ranking. Product pages establish feature and target-user contracts; measured systems are labeled separately.
Provider or maintainer says the capability exists.
Measured systemA defined human study exists, though often on an older system version.
Fastest route to a one-off result. Local semantics, delivery, and maintenance arrive mostly through the prompt and the person.
Executed analysis, tables, common charts, code, and downloads.
Control: inspect code and intermediate results.Generated HTML or application code for custom interactive explanation.
Watch: state, accessibility, hosting, maintenance.Specialist data chat for files, connections, charts, models, and reports.
Independent correctness and repair evidence is sparse.Installed context and handoff improve. Refresh behavior and copied data can quietly become the new failure surface.
Python-backed answers and optional static chart or table insertion.
Advanced mode can create editable, refreshable Python cells.Summaries, formulas, prompted edits, and editable inserted charts.
Inserted charts follow copied support data, not later changes to the original range.One workspace for spreadsheet data, live dashboards, slides, and team collaboration.
Evidence: provider contract.AI spreadsheet with Python, SQL, JavaScript, live connections, and visible code cells.
Control: schema and code remain in the grid.These environments give AI more project context while preserving cells, nodes, versions, and role-specific controls.
Edits SQL, Python, chart, pivot, and Markdown cells; builds apps; queries curated data.
Technical authors audit code; consumers use published apps; managers govern context.Adds tables, transformations, SQL, and charts beside existing work.
Control: creates new versions instead of silently overwriting or deleting.The assistant can make precise native edits because the editor constrains the available state.
Changes style, labels, sources, annotation, and accessibility settings in the editor.
Control: reversible native edits. Boundary: it does not edit source data or invent intent.More local context can narrow errors, but it also makes data teams responsible for metadata, measures, instructions, and review.
Creates and modifies report pages against model and report state.
Depends on measures, descriptions, permissions, and review.Builds views and calculations in existing web-authoring surfaces.
Control returns through the worksheet and data source.Maps questions onto governed LookML fields and measures.
Administrators own glossaries, defaults, and verified queries.Provider direction is converging on a separate author who owns semantics, permissions, review, and feedback.
Search and conversational analysis for business teams.
Data teams maintain governed models rather than answer queues.Shows sources, formulas, filters, and multi-step analysis.
Control: edit individual steps instead of regenerating everything.Structured and unstructured RAG plus chart and dashboard agents.
Includes permissions and audit logs.SQL-backed answers and visualizations in a curated no-code chat.
Authors monitor, review, and refine definitions and instructions.BI authoring, Q&A, executive summaries, and data stories.
Inherits datasets, topics, permissions, and author review.Charts and narratives for consumers over author-prepared data and metadata.
Explicitly separates author and consumer roles.These approaches make model output more inspectable and executable. Current repositories often exceed the versions evaluated in research.
Direct encoding plus natural-language transformation, branching, visible derived data, and reports.
Measured system Small reproduction study; current project is newer.Serializable declarative pipelines, charts, dashboards, and specialist agents.
Generated work can move from chat into notebooks or apps.Low-code Python dashboard specification with code escape hatches and agent tools.
Production components narrow generation; acceptance tests remain necessary.Retrievable library syntax, chart vocabulary, declarative output, and version guardrails.
Reports its own 174-case tests; independent current-model ablation is still needed.Creator experience
Public accounts and measured studies converge on a division of labor: coarse, repetitive, and reversible work benefits first; semantic decisions and precise refinement remain costly.
A first draft appears immediately. The idea becomes concrete enough to inspect and discuss.
The draft may encode an implicit aggregation, inherit a bad reference, or answer a nearby question.
AI: candidates and execution. Human: question, meaning, and selection.
Repetitive work collapses. SQL, calculations, docs, labels, bulk changes, and boilerplate move quickly.
Generated internals can be opaque; one strange implementation raises maintenance and explanation cost.
AI: legible repetition. Human: expected behavior and verification.
Many directions become cheap. A person can request alternatives without mastering every tool.
Filtering weak ideas requires expertise; novices may mistake breadth or polish for judgment.
AI: enumeration. Human: relevance, feasibility, and restraint.
A prototype improves the conversation. Stakeholders can react to layout and content before production data exists.
Prototype satisfaction says nothing about the correctness, refresh, permissions, or maintenance of the final system.
AI: negotiable mockup. Human: production contract and acceptance.
An unfamiliar representation becomes accessible. The model translates among prose, SQL, Python, tables, and charts.
A person without one familiar inspection surface may have no reliable way to evaluate the translation.
AI: translation. Human: check in a representation they understand.
Measured evidence
The studies below answer different questions. Their denominators and limits remain attached; the numbers should not be pooled into one score.
The exercise compared integrated ChatGPT analysis with an R/Stata-plus-ChatGPT workflow on simulated epidemiological data. Overall scores were not significantly different.
Twenty visualization novices completed 60 ChatGPT sessions producing 175 charts. The study observed prompting, chart quality, interpretation, verification, and repair—not just whether an image appeared.
Fatal task noncompliance52 of 60 sessions
Incorrect insight recorded12 of 60 sessions
Unusable chart22 of 175 charts
Meanwhile mean confidence was 3.73/5 and satisfaction 3.93/5. Only three verification attempts were observed.
Newer models produced fewer design flaws on replayed initial prompts, but richer interfaces added latency, broken controls, blank renders, and unverifiable interpretations. Better generation changed the failure surface; it did not remove the verification problem.
Use of explanation, code, original and intermediate data, results, and visual summaries
People began with procedure, then often moved to data after noticing trouble; expertise shaped the artifact they trusted
Prepared tasks at one company; data transformations rather than complete visualization projects
Mixed direct manipulation and natural language on reproduction tasks
All completed; visible transformed data, code, explanations, history, and branches supported different verification styles
Six needed hints; no baseline, open exploration, self-owned data, or long-term use
Rapid mockup generation, structured edits, and comparison with a lightly taught Tableau condition
Participants valued speed, simulated data, history, and a concrete object for negotiation
Pre-data prototypes, not analytical correctness, deployment, or ongoing dashboard use
AI versus human-expert design advice, ratings, interviews, and practitioner preference
AI helped enumerate ideas; practitioners preferred experts for accuracy, context, adaptability, and actionable advice
Practitioner sessions used 2023-era GPT-3.5; raw capability comparison is dated
At least five versions of a self-selected visualization over a three-to-five-day window with LLM and perceptual critique
Final work improved 3.69/5 on average; intermediate and expert designers converted feedback into useful edits more readily
Roughly 90–150 minutes of observed work each; no critique baseline, production delivery, or later maintenance
Correction, premature acceptance, non-completion, task time, hints, and perceived control across conversational and decomposed interfaces
Seven episodes were not completed and 31 were declared complete with an issue remaining; structure improved perceived control, not detected success or time
Tasks were engineered to contain model errors and stopped at 15 minutes; no publication, maintenance, or field abandonment
Analyst verification ↗ · Data Formulator 2 ↗ · DashChat ↗ · Visualization advice ↗ · Visualizationary ↗ · Steering and verification ↗
Human capability
A person can read a familiar dashboard but not code, know the domain but not visual design, or implement polished charts without being able to audit a misleading transformation. Evaluate the relevant capability, not one rank.
Read values, encodings, patterns, uncertainty, and unfamiliar forms.
Select data, choose a form, map fields, implement, annotate, and revise.
Test source fidelity, hidden transformations, misleading design, and omissions.
Relate the chart to domain meaning, audience, story, decision, and consequence.
Every competency is contextual. Data and statistical knowledge, domain semantics, visual design, implementation, situated judgment, and delivery experience are separate resources. AI may remove one barrier while leaving the others intact.
Growing the skill
That does not make direct work obsolete. It changes which difficulty deserves practice. The right test is what the person can explain, inspect, repair, and transfer after assistance is removed.
Tool access, syntax, debugging, scattered examples, and blank-page uncertainty kept many people from attempting the work.
Both mechanics and judgment required direct practice.Learners report faster coding and debugging. Proactive question-based scaffolding can improve immediate post-removal comprehension. Delayed construction transfer is mostly unmeasured.
Assisted performance and learning must be scored separately.More routine implementation will be delegated. Mental models, critique, verification, local repair, and reader responsibility remain the scarce work.
Revisit if answer-oriented assistance demonstrates delayed transfer to unfamiliar tasks.Framing · data semantics and statistics · critique · verification and calibration · alternative comparison · domain and audience judgment · accessibility · provenance and delivery
Direct construction · data wrangling · code and specification reading · sketching · hand-checking values · debugging transformations · precise local repair
API trivia · boilerplate · exhaustive taxonomy recall · manual pixel polishing · prompt incantations · deep recall of one tool's transient interface · first-render speed as a badge of skill
What “the hard way” should preserve: predict before revealing, translate questions into fields and encodings, check sample values, generate alternatives before seeing suggestions, diagnose before repair, explain decisions, and periodically transfer without assistance. Boilerplate and API hunting do not become educational merely because they are slow.
Semester-long visualization course study ↗ · Proactive scaffolding experiment ↗ · One-year visualization retention ↗ · Guardrails and unassisted learning ↗
Access to a first chart, code path, explanation, and more candidate ideas
Correctness, hidden-choice detection, verification, and reliable repair
Access: moderateAccepted work: low
Reported speed, engagement, confidence, mechanics reduction, and immediate post-removal comprehension from proactive scaffolding
Delayed independent construction and transfer to unfamiliar tasks; creativity and artifact-quality findings remain mixed or modest
Near transfer: promisingDelayed transfer: unknown
Turning critique, alternatives, and unfamiliar implementation into useful edits
A general “sweet spot”; the direct expertise-stratified sample is too small
PromisingNot settled
Bounded multiplication: option filtering, representation bridging, debugging, and constrained implementation
Open-ended judgment, production delivery, and a universal advantage over direct work
Bounded: moderateField: low
Translation of domain intent into a query, table, code candidate, or familiar chart
Audit of joins, measures, uncertainty, interaction, and generated implementation
Direct evidence: thin
Access is successful new work. Productivity is less total effort to an accepted artifact. Quality uses a declared correctness or usefulness rubric. Learning survives an unassisted transfer test. Verification detects and repairs defects. Reader outcome changes comprehension or decisions. Satisfaction and first-render speed do not stand in for the other rows.
One of eleven held fragments now matches a partial call/output-token budget, but its released accounting omits repair and failed-worker cost and its “final reports” are candidates. Zero comparisons join equivalent observed route-wide use to a frozen accepted-output denominator. Ask what was promised, what every attempt actually used, and how many became work you could accept.
The design requirement is not to pick one “user level.” Expose consequential choices for people with less construction skill, preserve precise control for experienced practitioners, and let every person verify through a representation they understand. One current experiment demonstrates immediate post-removal comprehension from proactive scaffolding; durable construction gain and AI-caused atrophy remain unestablished.
Visualization literacy review ↗ · Who counts as a novice? ↗ · Visualization learning ↗ · Creativity and time ↗ · Current novice study ↗ · Expert replication ↗
Reader experience
Creator delight is not a proxy for reader comprehension. The visual artifact inherits expectations from journalism, science, business, education, or social media.
Compression into marks, labels, annotations, interaction, prose, and provenance
→In a 2025 study, 31 of 37 participants mentioned clarity while ranking charts. Sources, integrity, familiar forms, and aesthetics also mattered, with substantial differences between people. The study measured deliberative trust, not correctness, comprehension, or behavior.
Trustworthy by Design ↗In a preregistered 2021 experiment, people brought strong preferences for human or algorithmic recommendations, but data relevance usually dominated actual selections. The same label suggested precision to some and missing human judgment to others.
Vis Ex Machina ↗On 60 synthetic charts, multimodal models enumerated data structure while 24 people formed trend narratives and reacted more to layout and overlap. Each could succeed on a different notion of intent.
How Do LLMs See Charts? ↗Across ten held rows, one directly carries selected AI-generated charts into an independent controlled-reader effect. Five measure adjacent human outcomes where AI labels, explains, assists, or sits beside a chart; three concern creators or co-designers; one uses VQA without people. Zero reaches accepted delivery and a later same-lineage reader recheck.
Read the eight-receipt boundary →In a randomized 117-person experiment, data stories, passive AI Q&A, and proactive scaffolded dialogue all improved visual-comprehension scores. After support was removed, the proactive group’s median was 6/6 versus 5/6 for both alternatives; completion time did not differ.
GenAI agents and visual comprehension ↗In a 48-person controlled experiment, second-phase accuracy was 71.9% for readers shown AI-generated misleading charts and 88.3% for readers shown correct charts. This measures bounded chart-QA harm—not prevalence, calibrated trust, or consequential decisions.
ChartAttack ↗A 26-person low-vision smartphone study found 100% completion with a full interactive treatment versus 61.5% with a screen magnifier. A separate 10-person blind-reader study found similar device-level accuracy but different time, workload, and chart-type results. Neither tested AI-generated charts.
Lexara follows six CVA developers using a deployed evaluation toolkit for two weeks across 38 experiments, 57 newly authored cases, ten models, and six prompts. They made and challenged model/prompt selections using their own data, but the study does not join an immutable build and exact configuration to accepted downstream delivery, an intended-reader decision or calibrated trust, whole cost, and a later post-release recheck. Keep development selection, contract acceptance, audience consequence, confidence calibration, and return use separate. Read the twelve-receipt account →
One explicit GPT study among 122 stable-key titles adds a pin-able, 1,271-blob supplement of prompts, results, generated code, and grader bundles. It does not add an immutable provider snapshot, versioned release, accepted final project, delivery to the government, UN, or student audiences named in prompts, an audience outcome, or a later same-lineage return. Treat the other 121 title nonmatches as not surfaced, not excluded. Read the complete custody boundary →
AIDSVu adds aggregate use, named planning applications, recurring governance, explicit public-data states, and a later 2026 data release. Those are mature platform receipts. The held surfaces expose no AI visualization role, immutable build, version-bound measured audience outcome, affected-audience recheck, or whole cost. Keep it as a non-AI comparator and keep 86 abstract cue nonmatches unexcluded. Read the complete afterlife boundary →
Five recovered full texts separate controlled audience measurement, public-sector co-design/demo/bounded use, and enterprise demonstrator plus production intent. None binds an accepted field release to a later affected-audience recheck, and all five full-text GenAI cue screens are empty. Later recovery adds 20 substantive surfaces, including one further full chapter, and leaves two DOI rows content-unassessed. Co-design, a demo, a human task study, and production intent are different receipts—not deployment. Read the complete receipt ladder →
Eight stable non-DOI keys yield three DOI repairs, two year corrections, one rejected foreign PMID, and four full texts. The EMR cancer diary reports increasing system-log use; an independent review recovers an 11-clinician QUIS median of 4.38. That is meaningful embedded-use and usability evidence, but it supplies no AI role, immutable accepted build, patient or decision outcome, later event, or affected-clinician recheck. One primary text remains gated and unassessed. A field-use receipt is still not an AI lifecycle. Read the complete field-use boundary →
Nine high-signal DOI rows yield eight substantive primary abstract or official-project surfaces and seven participant/evaluator studies. InfoViP is the delivery near-miss: seven FDA safety evaluators shaped and evaluated the prototype, suggestions were addressed, and an official page says an enhanced NLP and unsupervised-learning version will be installed in production. Future tense is not installation, acceptance, routine use, regulatory outcome, later change, or evaluator return. Subsequent recovery reduces the content-unassessed DOI remainder from 14 to two. Read the complete human-evaluation boundary →
The final 2025 CIOMS report supports an approved AWS/AERS component processing more than 30 million historical plus about 8,000 daily submissions while leaving a solid QA plan, completed audits, routine roles, and signal effects open. The now-closed Elsa/API/UI opportunity names Joshua Xu, Leihong Wu, and Oanh Dang as research mentors, but supplies no selection, work, acceptance, or release receipt. A fellow is a nonemployee barred from inherently governmental functions. Neither the actor profiles nor an adjacent AI-QA project assigns InfoViP operation, maintenance, QA execution, release approval, authorization, a completed audit, or the Elsa application join. Read the actor-and-authority boundary →
A full chapter connects three UX experts and 25 distinct problems to an implemented third version. A corridor stakeholder case and university-network case add practice context. None exposes a GenAI visualization role, accepted field release, routine-use denominator, consequential outcome, or later affected-actor return. Expert-evaluated, redesigned, and applied with stakeholders are useful stages—not synonyms for deployment. Read the complete residual account →
B110 now has an official abstract, pinned framework source, a pinned DiscoverWater application, and a same-named KU interface reachable in August 2026. At that pass, B92 and B91 remained content-unassessed. The live bytes are not bound to a SHA; acceptance, continuous or ordinary use, audience outcome, recheck, whole cost, and an AI role remain absent from held surfaces. Read the complete delivery boundary →
The dated live page differs from the sole published v1.2 page state; the linked v2.0 implementation is R/Shiny. Eight of twelve dependency names occur in pinned application source, and two of three sampled data assets match canonically. The partial lineage is real, but no manifest binds the deployed page to source. A prototype demonstration plus analytics and comment hooks add no acceptance, use, audience outcome, accessibility, later-return, whole-cost, or AI receipt. Read the complete source-lineage boundary →
B92's exact publisher abstract describes multicriteria risk evaluation, Monte Carlo simulation, Kendall's tau rank comparison, and graphs for ordering pipeline sections. B91 remains content-unassessed after three bounded lawful passes over six named surfaces. The content ledger stays 6 full / 20 abstract or official / 1 unassessed / 0 complete AI lifecycles; the separate workflow state is zero active generic B91 targets. Reopen only when exact new lawful custody appears. Paused is not negative. Read the complete stop-and-reopen boundary →
The complete register binds 122 of 127 supplement positions to distinct A208 review keys. Five exact-looking DOI identities have no A208-controlled join, and one admitted key carries a PubMed identifier that points to a different paper. Keep source label, review key, candidate, identifier validation, and admission authority separate. Lifecycle screening is unstarted. Never promote the five tempting matches, reuse the conflicted PMID, score an unresolved row, or report zero of 127. Read the complete boundary →
Context
Preserve the question, data, decisions, artifact, provenance, and acceptance evidence. Then match the authoring and evaluation method to the human purpose.
Helpful: exploration, annotation alternatives, accessible implementation, source checks.
Danger: fluent generation substituting for reporting, authorial judgment, or reader testing.
Helpful: governed measures, stable layout edits, anomaly explanation tied to source visuals.
Danger: silent changes to metrics, thresholds, state, or hierarchy.
Helpful: semantic grounding, visible intermediate data, reusable verification.
Danger: confident answers over weak metadata, ambiguous measures, or hidden filters.
Helpful: many cheap views, branches, undo, direct manipulation.
Danger: turning the first plausible pattern into the final narrative.
Helpful: restricted representations, executable transformations, linked views, domain checks.
Danger: visual plausibility standing in for scientific correctness.
Helpful: adjustable assumptions, guided interaction, feedback, multiple representations.
Danger: generating interactivity without measuring what learners understand.
The lifecycle test
Question → data authority → generation → checking → correction → delivery → reader use → sustainment. The spine stays fixed; acceptance changes with the environment.
Current diaries expose source work, decomposition, defects, correction, mobile checks, publishing, and proposed updates.
Independent verification, editorial review, reader comprehension, accessibility, later refreshes, and civic use.
Source-to-claim trace, editorial gate, delivered desktop/mobile/access states, audience test, correction policy, and update owner.
Product contracts expose semantic models, queries, permissions, and review surfaces; testimony explains the need for narrow, owned metrics.
Routine task success, review queues, refusal, one-number incidents, trust recovery, and total cost in one independent study.
Approved metric contract, exact query and filters, permissions, owner, refusal behavior, escalation, and decision outcome.
Three same-artifact cases cross public delivery into a later event; two reach accepted repair, corrected delivery, and maintainer recheck.
Affected-user or independent-operator recovery is 0/3, transferred authority is 0/3, and complete twelve-state rows are 0/3. Whole cost, accessible-reader use, decision outcome, and calibrated trust also remain missing.
Keep provenance, release, exposure, event, diagnosis, repair, corrected delivery, maintainer recheck, affected-actor recheck, prevention, contribution, and authority separate.
One 117-person experiment measures immediate comprehension; one classroom case shows direct instruction rescuing failed AI-assisted construction.
Delayed construction, unfamiliar transfer, ordinary assistant use, and a complete learner-to-reader project.
Immediate and delayed unassisted performance, diagnosis, transfer, confidence calibration, and the eventual audience’s comprehension.
Controlled studies measure low-vision smartphone and blind nonvisual outcomes. MAIDR adds pre-AI and abstract-level AI study evidence plus a version floor. Graphy adds the same three blind co-designers returning for 12 sessions across four workshops and eight months while the interaction changed.
MAIDR's tested build and model remain unknown. Graphy is not a formal usability or performance evaluation; its workshops have no immutable code/model binding, and its seven-commit repository has no tag, release, or representative recheck.
Keep co-design, formal evaluation, study build, first containing release, changed release, recheck, and rewrite separate; compare exact AI versions with representative users on their own assistive technology.
Put the evidence to work
Three same-artifact cases join an AI-assisted dashboard to a later event, and two reach maintainer-verified restoration. None receives an affected-actor recovery receipt or transfers maintenance authority. No captured episode also joins whole cost, accessible reader use, a consequential decision, and calibrated trust.
Preserve AI provenance, pre-event delivery, actor-separated exposure, later event, diagnosis, accepted repair, corrected delivery, maintainer recheck, affected-user or independent-operator recheck, prevention, accepted second-person change, and transferred authority. Zero case clears all twelve. OpenClaw issue 30 and PR 31 remain open; Prism issue 81 is closed after a maintainer check but without the affected operator's return. Ticket status is not a recovery receipt. Read the complete account →
The CHI 2024 MAIDR study gives 11 blind participants a pre-AI multimodal chart system; the current project declares a complete TypeScript rewrite and an AI description layer. That is a valuable baseline plus a new artifact state—not an exact-version participant recheck. A public-health copilot paper describes a 16-person trust/usability method, then says the experiment was removed and reports no human result. Join exact artifact and version, AI state, release, later event, actor, custody, and outcome before counting a lifecycle.
Public code narrows the later eight-participant AI study to a legacy box-plot surface and v2.10.0 as the first containing tag. It does not name the participant-tested build or model. Later model swaps, verification fixes, deprecation, and the separate TypeScript rewrite trigger new representative checks; they do not supply them.
Graphy follows the same three blind co-designers while selected interaction changes are implemented between rounds. The final own-data workshop shows the evolved interaction in use, but the paper explicitly is not a formal usability or performance evaluation. Its seven-commit public repository has no workshop-version binding, tag, release, or representative recheck. Count repeated co-design—not a release afterlife.
Every required receipt appears somewhere in the held evidence, but none stays attached to one artifact lineage. A complete row needs repeated representative use, immutable tested build, exact exposure model, versioned release, later material event, and representative or actor-separated post-change recheck. This is zero of seven in named surfaces through 15 August 2026—not a claim about every private or inaccessible case.
Three returning blind co-designers across 12 sessions and eight months.
Workshop build and exact model; no release.
Bind the workshop, then recheck a changed release.
BLV study, study surface, version floor, and later maintenance.
Participant-tested build and model.
Return-user check after a pinned maintenance change.
Repeated repair attempts, v1.8.11, maintainer runtime check.
Exact model and affected-operator final recheck.
Operator exercises the corrected release.
Tagged release, non-owner failure, open repair proposal.
Corrected release and reporter recheck.
Merge → release → reporter reconciliation.
Live regression, repair, restored production, prevention.
Exact model and actor-separated recheck.
Independent user or operator recovery check.
Frozen v0.7.4 ledger: 18 sessions, 151 commits, 20 releases.
Representative use and elapsed field event.
Recipient returns after a released change.
50 native repairs and 80 changed visualizations.
Delivered artifact and returning user.
Carry one accepted output into real afterlife.
A 2021–2025 HealthTech program followed 21 projects through 84 recorded calls, notes and decision logs, backlogs and issue trackers, one governance-dashboard case, and a 16-startup survey. The case reached deployment with patients. It is longitudinal visualization evidence—not a generative-authoring study—and it does not measure total cost, accessible reader outcomes, calibrated trust, or comparative maintenance efficacy.
DV-World adds 50 native Excel repair tasks and 80 new-data or evolving-requirement tasks. The best reported agents reached 48% repair success and 51.44% evolution. Use a prepared change as a pre-delivery gate; the benchmark does not follow an accepted artifact through elapsed maintenance, handoff, reader use, or whole cost.
A four-month KubeStellar Console report and commit-pinned QA record follow one AI-assisted dashboard through a live blank-page regression, two-step repair, restored deployment, and a manually added post-build check. This is project-authored single-maintainer evidence—not a comparative rate, whole-cost ledger, handoff, accessible-reader result, consequential decision, or trust result.
OpenClaw Agent Dashboard says it was built with Claude Code and shipped v3.0.0. Thirteen days later, a non-owner operator reported that its cost view showed $0 on a custom provider despite present token data. The owner acknowledged that configuration had not been tested. A different non-owner's open PR 31 proposes a matching fallback, but has no maintainer review, merge, corrected release, or reporter recheck. A proposal is not recovery.
Prism says its family dashboard was built with Claude Code under human product direction. One non-owner operator's Home Assistant route failed at install, repair build, and later startup before they used ordinary Docker instead. v1.8.11 and its entrypoint contain the named socket and schema-replay repairs; the maintainer reports install, start, and restart checks on real Home Assistant OS. The operator never rechecked it. Route substitution is not recovery of the failed route, and no defect is attributed to Claude Code.
Prism's non-owner PR 23 carries a Claude-assisted weather visualization through owner-found defects, contributor repair, owner integration, and a credited release; the contributor later returns with accepted PR 41. OpenClaw's non-owner PR 15 puts Claude-assisted multi-provider pricing into the product after owner integration. These establish second-person change—and one repeat contributor—not release, incident, or recovery authority.
Non-owner changes the exact artifact; owner accepts it.
Prism PRs 23/41; OpenClaw PR 15.
Ownership or incident duty.
Same non-owner returns after elapsed time.
Prism, once after 17 days.
Durable team membership.
Named second person can release or respond and exercises that role.
Missing.
A contributor is a maintainer.
Affected or separate operator exercises the corrected route.
Missing in Prism issue 81 and OpenClaw issue 30.
Open patch or owner check is recovery.
Project-authored live failure, two-step repair, restored deploy, prevention check.
Independent acceptance, whole cost, handoff, accessible readers, decisions, trust.
Keep fix, deploy, live recovery, and future gate as separate receipts.
Non-owner report plus a different non-owner's open repair proposal.
Maintainer acceptance, merge, corrected release, reporter recheck, invoice reconciliation, downstream outcome.
Proposal is not accepted repair; close only after release and independent recheck.
Non-owner install, repair-build, and startup failures; corrected release; maintainer runtime check.
Affected-operator recheck, whole cost, handoff, accessible readers, decisions, trust.
Keep report → repair → release → maintainer recheck → independent recheck as five receipts.
Intended audience, decision, source authority, local definitions, baseline.
Can the creator explain the claim? Who owns the metric, judgment, and consequences?
Were task, stakes, expertise, and current comparison declared before use?
Prompts, transformations, direct edits, failures, waits, rollback, rejected output.
What was inspected or repaired? What review changed the artifact, and what stayed disputed?
Measure time to first candidate separately from time to accepted work.
Named approver, acceptance contract, published state, desktop, mobile, keyboard, and assistive-technology checks.
Did the real surface pass? Keep approval, publication, and reader use as separate states.
Recompute claims and capture the delivered state independently.
Defined reader tasks, comprehension, decisions, confidence, errors, feedback.
What did intended readers understand and do? Who was excluded?
Measure decision quality, calibration, recovery, and accessibility—not satisfaction alone.
Refresh host, dependencies, second maintainer, incidents, correction policy, retirement, human and model cost.
Can someone else update or retire it? Who owns the next refresh and failure response?
Follow a dependency change or handoff; compare whole cost against current direct work.
A practical evaluation
A prettier first render is not enough evidence for a workflow change.
Exploration, governed BI, operations, public explanation, science, education, or a reusable application carry different obligations.
Include self-owned data, awkward semantics, a correction request, and the actual delivery surface. Retain a simple task as a control.
Measure human time, errors, correction path, and delivery effort before claiming a gain.
Count prompting, waiting, cleanup, verification, publishing, and monetary cost.
Make fields, filters, aggregation, transformations, code or semantic query, and generated interaction state visible.
Recompute values, exercise controls, check exports, accessibility, responsive behavior, and future maintenance.
Ask defined readers for the claim, evidence, uncertainty, and next action. Measure correctness, time, confidence, and harmful misreadings.
The updated gap ledger
“Partly answered” means one bounded study exists—not that the field can generalize.
How to read the evidence
Supports claims about its participants, tasks, models, measures, and comparisons. Small or dated studies do not automatically transfer to current production.
Shows which jobs and tools respondents report. Self-selection, missing items, and sponsor or community channels limit population claims.
Shows what the provider says the feature can see and do. It does not independently establish correctness, usability, or adoption.
Surfaces jobs and failure costs that experiments may omit. Self-selection means it cannot establish prevalence or comparative performance.
Evidence cut 16 August 2026. Twenty-three human studies and structured workplace evaluations and one 127-study systematic review were read in full alongside two current single-case AI-assisted development reports, one independently reported open-source dashboard failure chain, practitioner surveys, current provider and open-source documentation, public discussions, first-person production cases, and dated essays and interviews from several professional positions. The strongest counter-reading is that better models will erase older failures. Fresh evidence shows real improvement in visual output—and new failure surfaces from richer, slower, more complex artifacts. Capability is moving; the location of the human work is moving with it.