Can it make the thing?
A chart renders. A dashboard runs. A prompt produces an analysis or interactive artifact.
Evidence: execution, task completion, feature contractResearch snapshot · August 2026
AI is already good at getting someone from a blank page to a plausible first thing. It is much less reliable at carrying the work through local meaning, correction, delivery, and audience understanding.
The reliable win is compression, not delegation. Faster creation matters; it does not transfer responsibility for the analytical claim or the reader’s understanding.
For creators, readers, and product teamsThis companion translates research results into the lived jobs, gains, costs, and trust questions surrounding current tools.
Not a product rankingProducts appear to explain interaction environments. Documentation establishes features, not comparative quality.
The central gap
Most demos answer the first question below. Practitioners live in the second. Readers determine the third.
A chart renders. A dashboard runs. A prompt produces an analysis or interactive artifact.
Evidence: execution, task completion, feature contractThe data and semantics are right. Decisions are inspectable. Precise edits, delivery, and reuse are possible.
Evidence: accepted artifact, total time, errors, repair, maintenanceThe creator can explain the process. The reader grasps the claim, limits, and appropriate next action.
Evidence: observed behavior, reader tasks, decisions, confidence calibrationWho is reaching for what
The best available adoption evidence shows practitioners adding AI beside spreadsheets, code, design, and BI tools. Creation and analysis are ahead of trusted conversational consumption.
Online, self-selected survey. The result describes respondents, not the population or product market share.
Use was up 13 percentage points from the 2023 survey. The 28 unsure responses are the narrow violet segment.
Multiple selections were allowed, so the bars do not sum to 100%.
A general assistant beside Excel, Tableau, Python, R, D3, Figma, or a publishing tool
Prepare, debug, learn, storyboard, draft labels, make a first view
Observed survey patternSpecific AI-product share was not measured.
General coding assistance, then project-aware notebook or development agents
Write SQL or Python, document models, debug pipelines, create analysis output
Observed survey patternBroader than visualization and vendor-adjacent.
AI inside Excel or Sheets; AI spreadsheets when code or live connections outgrow the grid
Ask about a range, create formulas and charts, preserve a familiar handoff
Installed-base inferenceExcel use is observed; AI-feature adoption is not.
AI inside an existing BI, notebook, or governed data platform
Author reports, reuse measures, inspect queries, answer follow-ups, govern access
Provider target + casesNo independent head-to-head field test.
Conversational BI over curated metrics
Retrieve a number, ask why it changed, get an ad hoc cut without navigating a report
Provider target + testimonyTrust depends on narrow, owned definitions.
General help for ideation or code, then a deliberate design and publishing surface
Explore, create variants, annotate, improve accessibility, implement a story
Small subgroup + inferenceNot a population estimate.
DVS survey and public data ↗ · Analytics-engineering survey ↗. Both are self-reported and nonrepresentative; denominators remain attached.
The jobs
“Make a chart” hides work before, during, and after visual encoding. The same assistant can be strong at one job and harmful at the next.
Decide what is worth doing and which data is allowed to answer it.
Turn source material into evidence without losing its meaning.
Choose what the audience will see and make every consequential choice inspectable.
Test the real surface, support the reader, and keep the work correct later.
Listening to the field
The same person can be delighted by a first draft, frustrated by the repair loop, afraid to put their name on the result, and still choose to use the tool tomorrow. Attitude and behavior do not move together.
People arrive with organizational demands, prior skill, career hopes, and fears about what assistance may remove.
Leadership wanted visible thought leadership. The custom chatbot was unreliable; for the immediate task, a pivot table was faster. The frustration was having to perform adoption.
Jisell Howe first worried that instant code would remove the journey from idea to customized chart. A concrete build changed the feeling: errors remained, but search and troubleshooting became less disruptive.
AI produced a first bar chart, but an instructor supplied the visualization knowledge needed to improve it. The gain was reaching the work—not yet independent mastery.
Delight tends to attach to access, flow, and continuity across chores, not only to a chart appearing.
A two-day revision felt surprisingly fluid. Four short lines were enough once the current files, requested changes, and source URLs were present. The creator’s lesson was that organized context mattered more than prompt polish.
Sef Kloninger called the work “plain fun.” The agent moved with him from a 500 GB dataset through tests, exploration, a dashboard, debugging, and a presentation; he still inspected raw data and requested checks.
The practitioner immediately added the caveat: prior proficiency made the compression possible. A novice in another account valued a smaller gain—producing cleaning and validation scripts that had been out of reach.
The experience depends on whether explaining, waiting, inspecting, and retrying costs less than direct manipulation.
Stale filters, bookmark identifiers, whitespace churn, and slow edit-preview cycles accumulated. One experienced practitioner said describing a small change could take longer than making it.
Christina Stathopoulos described useful brainstorming and exploration, then recalled a generated chart whose accompanying interpretation reversed the visible comparison. The feared failure was looking careless in front of a stakeholder.
A recognizable site appeared in about an hour. Geography, evidence notes, rate limits, mobile behavior, architecture, and fact-checking then consumed repeated rounds of repair.
The imagined stakeholder, patient, client, colleague, or excluded reader changes which errors matter and where refusal is rational.
An agent recreated roughly a week’s visible dashboard work from a finished screenshot and source tables. The consultant’s question was what clients had really paid for: construction, diagnosis, or the judgment embedded in the reference.
Participants welcomed boilerplate, inspiration, and translation while some rejected final AI imagery for anatomy, patient communication, or scientific work. Others protected rendering because it was also where control, flow, and creative joy lived.
Johny Cassidy describes exclusion from charts with missing or useless descriptions—and the relief of colleagues taking responsibility. Generated alt text is not accessible delivery without alternatives, feedback, and correction.
Use, maintenance, correction, learning, and second-person handoff determine whether the saved effort became value.
A manager praised an urgent dashboard and sent it to six colleagues. Three months later, the usage record showed one manager view. Creator pride, stakeholder approval, and reader use were three different outcomes.
AI documented SQL and calculations, reviewed junior work, and rehearsed stakeholder questions. Those tasks made the dashboard easier to explain and maintain without asking the model to own the final claim.
Emilia Ruzicka collected and drew personal data by hand. Slowness, imperfection, and direct contact with the data produced experimentation and care—the kind of learning a faster route can accidentally remove.
Listening past the original creator
A second pass pursued spreadsheet-native work, people who do not make AI their default, recipients deciding whether an answer is defensible, blind and low-vision learners, and the second person asked to maintain the result.
DWP staff allocated tasks according to expertise, time, trust, data sensitivity, and habit. Data-heavy Excel work, chart generation, and intricate formatting remained weak points. One person valued having the assistant but was often too busy to remember to use it.
Chat 15Dashboard 3Both 2
Chat 0Dashboard 18Both 2
A supply-chain analyst noticed that AI had displaced visits to a peer forum. Replies described both useful help and invented functions, wrong references, damaged formulas, and refusal. Several people valued the incidental learning produced by solving someone else’s problem in public.
Program managers and funders wanted overview, drill-down, definitions, neutral language, and review support. In a separate sales-dashboard demonstration, an observer rejected the idea that an executive seeking a quick answer should independently validate regenerated charts against raw data.
Eleven participants preferred tactile charts plus text and chat; one said the better mode depended on complexity; none preferred text and chat alone. Touch supplied spatial structure and chat supplied flexible clarification. Measured chart-understanding accuracy did not improve.
An AI-built dashboard failed while its creator was away. The inheriting maintainer replaced the creator-local scheduled job with a proper pipeline; replies surfaced missing metric semantics, hidden dependencies, documentation gaps, and service expectations nobody had owned.
What listening adds
Field perspectives
These sources expose purposes and production constraints that benchmarks omit. Most are first-person cases, interviews, or essays: they establish that a perspective or experience exists, not how common or effective it is.
Two current newsroom diaries describe work that once took weeks appearing in days or hours.
They also contain wrong totals, broken geography, missed notes, repeated correction, usage limits, mobile checks, and new validation work. Benn Stancil explains why a plausible chart cannot validate its calculation.
Semantic stakes, independence of the check, cost of error, and ownership of the complete pipeline.
Enrico Bertini maps opportunities across data acquisition, wrangling, analysis, tours, and creative exploration. Alberto Cairo treats easier code as capacity extension.
Question quality, interpretation, and final analytical choice become more important. Exploratory, explanatory, essayistic, and artistic work do not share one quality function.
Purpose, audience, consequence, capability profile, and whether the claim remains inspectable.
Removing syntax and formatting work can get a person to a candidate before the question goes cold.
A classroom exercise used failed prompting to reveal missing chart structure. An analog practice made slowness a source of attention, experimentation, and care.
Whether the effort is obsolete mechanics or practice in structure, skepticism, memory, and meaning.
Generated descriptions and alternatives may help newsrooms address access at a scale that manual workflows have not reached.
Elavsky and Xiong Bearfield show model chains losing data, source, uncertainty, design, and purpose. Blind journalist Johny Cassidy describes an organizational and multimodal problem, not an alt-text checkbox.
Whether the reader can test the account, choose another modality, report failure, and obtain remediation.
Richard Brath asks which charts remain useful when models extract insights. Elijah Meeks proposes audience, intent, interpretation, and conversation as framework objects.
Domestic Data Streamers argues that visualization remains a distinct language for comparison, human experience, and creative divergence.
Reader task, need for overview or shared evidence, qualitative meaning, novelty, and an auditable interaction trace.
Production cases
Both are self-reports by the people who built and checked their own projects. Keeping the gain beside the incurred work prevents “built in a week” from becoming a quality claim.
One map fell from roughly three weeks in 2012 to a two-day reconstruction.
A recognizable site appeared in about an hour; the full dashboard and checking workflow took one week.
An equal-budget comparison with the best current direct workflow or another practitioner.
Incorrect totals and incomplete bilingual text required source comparison and repeated correction.
Cities rendered in the sea or broken shapes; exact evidence and source notes were sometimes missed.
A field distribution of error, correction time, regression, premature acceptance, or abandonment.
About 20 ordered tasks, source checks, and open code made components easier to trace and revisit.
Extraction, data, presentation, insights, and fact-checking were separated so repairs could be replayed.
That decomposition or a model-built validator independently improves correctness across projects.
Working prototypes were published; attention, understanding, and civic use remained the stated bottleneck.
Mobile rendering was checked and stakeholders reacted; readers were not tested.
Audience understanding, accessibility, consequential use, or sustained adoption.
Record reconciliation and scheduled updates were proposed next steps.
Refactoring and longer-term maintenance remained after the one-week build.
Later refreshes, dependency changes, second-person handoff, incident response, or retirement.
Ottaviani build diary ↗ · Tordecilla build diary ↗. These cases support process claims, not general productivity or accuracy estimates.
The tools
This is a showcase of current approaches, not a ranking. Product pages establish feature and target-user contracts; measured systems are labeled separately.
Provider or maintainer says the capability exists.
Measured systemA defined human study exists, though often on an older system version.
Fastest route to a one-off result. Local semantics, delivery, and maintenance arrive mostly through the prompt and the person.
Executed analysis, tables, common charts, code, and downloads.
Control: inspect code and intermediate results.Generated HTML or application code for custom interactive explanation.
Watch: state, accessibility, hosting, maintenance.Specialist data chat for files, connections, charts, models, and reports.
Independent correctness and repair evidence is sparse.Installed context and handoff improve. Refresh behavior and copied data can quietly become the new failure surface.
Python-backed answers and optional static chart or table insertion.
Advanced mode can create editable, refreshable Python cells.Summaries, formulas, prompted edits, and editable inserted charts.
Inserted charts follow copied support data, not later changes to the original range.One workspace for spreadsheet data, live dashboards, slides, and team collaboration.
Evidence: provider contract.AI spreadsheet with Python, SQL, JavaScript, live connections, and visible code cells.
Control: schema and code remain in the grid.These environments give AI more project context while preserving cells, nodes, versions, and role-specific controls.
Edits SQL, Python, chart, pivot, and Markdown cells; builds apps; queries curated data.
Technical authors audit code; consumers use published apps; managers govern context.Adds tables, transformations, SQL, and charts beside existing work.
Control: creates new versions instead of silently overwriting or deleting.The assistant can make precise native edits because the editor constrains the available state.
Changes style, labels, sources, annotation, and accessibility settings in the editor.
Control: reversible native edits. Boundary: it does not edit source data or invent intent.More local context can narrow errors, but it also makes data teams responsible for metadata, measures, instructions, and review.
Creates and modifies report pages against model and report state.
Depends on measures, descriptions, permissions, and review.Builds views and calculations in existing web-authoring surfaces.
Control returns through the worksheet and data source.Maps questions onto governed LookML fields and measures.
Administrators own glossaries, defaults, and verified queries.Provider direction is converging on a separate author who owns semantics, permissions, review, and feedback.
Search and conversational analysis for business teams.
Data teams maintain governed models rather than answer queues.Shows sources, formulas, filters, and multi-step analysis.
Control: edit individual steps instead of regenerating everything.Structured and unstructured RAG plus chart and dashboard agents.
Includes permissions and audit logs.SQL-backed answers and visualizations in a curated no-code chat.
Authors monitor, review, and refine definitions and instructions.BI authoring, Q&A, executive summaries, and data stories.
Inherits datasets, topics, permissions, and author review.Charts and narratives for consumers over author-prepared data and metadata.
Explicitly separates author and consumer roles.These approaches make model output more inspectable and executable. Current repositories often exceed the versions evaluated in research.
Direct encoding plus natural-language transformation, branching, visible derived data, and reports.
Measured system Small reproduction study; current project is newer.Serializable declarative pipelines, charts, dashboards, and specialist agents.
Generated work can move from chat into notebooks or apps.Low-code Python dashboard specification with code escape hatches and agent tools.
Production components narrow generation; acceptance tests remain necessary.Retrievable library syntax, chart vocabulary, declarative output, and version guardrails.
Reports its own 174-case tests; independent current-model ablation is still needed.Creator experience
Public accounts and measured studies converge on a division of labor: coarse, repetitive, and reversible work benefits first; semantic decisions and precise refinement remain costly.
A first draft appears immediately. The idea becomes concrete enough to inspect and discuss.
The draft may encode an implicit aggregation, inherit a bad reference, or answer a nearby question.
AI: candidates and execution. Human: question, meaning, and selection.
Repetitive work collapses. SQL, calculations, docs, labels, bulk changes, and boilerplate move quickly.
Generated internals can be opaque; one strange implementation raises maintenance and explanation cost.
AI: legible repetition. Human: expected behavior and verification.
Many directions become cheap. A person can request alternatives without mastering every tool.
Filtering weak ideas requires expertise; novices may mistake breadth or polish for judgment.
AI: enumeration. Human: relevance, feasibility, and restraint.
A prototype improves the conversation. Stakeholders can react to layout and content before production data exists.
Prototype satisfaction says nothing about the correctness, refresh, permissions, or maintenance of the final system.
AI: negotiable mockup. Human: production contract and acceptance.
An unfamiliar representation becomes accessible. The model translates among prose, SQL, Python, tables, and charts.
A person without one familiar inspection surface may have no reliable way to evaluate the translation.
AI: translation. Human: check in a representation they understand.
Measured evidence
The studies below answer different questions. Their denominators and limits remain attached; the numbers should not be pooled into one score.
The exercise compared integrated ChatGPT analysis with an R/Stata-plus-ChatGPT workflow on simulated epidemiological data. Overall scores were not significantly different.
Twenty visualization novices completed 60 ChatGPT sessions producing 175 charts. The study observed prompting, chart quality, interpretation, verification, and repair—not just whether an image appeared.
Fatal task noncompliance52 of 60 sessions
Incorrect insight recorded12 of 60 sessions
Unusable chart22 of 175 charts
Meanwhile mean confidence was 3.73/5 and satisfaction 3.93/5. Only three verification attempts were observed.
Newer models produced fewer design flaws on replayed initial prompts, but richer interfaces added latency, broken controls, blank renders, and unverifiable interpretations. Better generation changed the failure surface; it did not remove the verification problem.
Use of explanation, code, original and intermediate data, results, and visual summaries
People began with procedure, then often moved to data after noticing trouble; expertise shaped the artifact they trusted
Prepared tasks at one company; data transformations rather than complete visualization projects
Mixed direct manipulation and natural language on reproduction tasks
All completed; visible transformed data, code, explanations, history, and branches supported different verification styles
Six needed hints; no baseline, open exploration, self-owned data, or long-term use
Rapid mockup generation, structured edits, and comparison with a lightly taught Tableau condition
Participants valued speed, simulated data, history, and a concrete object for negotiation
Pre-data prototypes, not analytical correctness, deployment, or ongoing dashboard use
AI versus human-expert design advice, ratings, interviews, and practitioner preference
AI helped enumerate ideas; practitioners preferred experts for accuracy, context, adaptability, and actionable advice
Practitioner sessions used 2023-era GPT-3.5; raw capability comparison is dated
At least five versions of a self-selected visualization over a three-to-five-day window with LLM and perceptual critique
Final work improved 3.69/5 on average; intermediate and expert designers converted feedback into useful edits more readily
Roughly 90–150 minutes of observed work each; no critique baseline, production delivery, or later maintenance
Correction, premature acceptance, non-completion, task time, hints, and perceived control across conversational and decomposed interfaces
Seven episodes were not completed and 31 were declared complete with an issue remaining; structure improved perceived control, not detected success or time
Tasks were engineered to contain model errors and stopped at 15 minutes; no publication, maintenance, or field abandonment
Analyst verification ↗ · Data Formulator 2 ↗ · DashChat ↗ · Visualization advice ↗ · Visualizationary ↗ · Steering and verification ↗
Human capability
A person can read a familiar dashboard but not code, know the domain but not visual design, or implement polished charts without being able to audit a misleading transformation. Evaluate the relevant capability, not one rank.
Read values, encodings, patterns, uncertainty, and unfamiliar forms.
Select data, choose a form, map fields, implement, annotate, and revise.
Test source fidelity, hidden transformations, misleading design, and omissions.
Relate the chart to domain meaning, audience, story, decision, and consequence.
Every competency is contextual. Data and statistical knowledge, domain semantics, visual design, implementation, situated judgment, and delivery experience are separate resources. AI may remove one barrier while leaving the others intact.
Growing the skill
That does not make direct work obsolete. It changes which difficulty deserves practice. The right test is what the person can explain, inspect, repair, and transfer after assistance is removed.
Tool access, syntax, debugging, scattered examples, and blank-page uncertainty kept many people from attempting the work.
Both mechanics and judgment required direct practice.Learners report faster coding and debugging. Proactive question-based scaffolding can improve immediate post-removal comprehension. Delayed construction transfer is mostly unmeasured.
Assisted performance and learning must be scored separately.More routine implementation will be delegated. Mental models, critique, verification, local repair, and reader responsibility remain the scarce work.
Revisit if answer-oriented assistance demonstrates delayed transfer to unfamiliar tasks.Framing · data semantics and statistics · critique · verification and calibration · alternative comparison · domain and audience judgment · accessibility · provenance and delivery
Direct construction · data wrangling · code and specification reading · sketching · hand-checking values · debugging transformations · precise local repair
API trivia · boilerplate · exhaustive taxonomy recall · manual pixel polishing · prompt incantations · deep recall of one tool's transient interface · first-render speed as a badge of skill
What “the hard way” should preserve: predict before revealing, translate questions into fields and encodings, check sample values, generate alternatives before seeing suggestions, diagnose before repair, explain decisions, and periodically transfer without assistance. Boilerplate and API hunting do not become educational merely because they are slow.
Semester-long visualization course study ↗ · Proactive scaffolding experiment ↗ · One-year visualization retention ↗ · Guardrails and unassisted learning ↗
Access to a first chart, code path, explanation, and more candidate ideas
Correctness, hidden-choice detection, verification, and reliable repair
Access: moderateAccepted work: low
Reported speed, engagement, confidence, mechanics reduction, and immediate post-removal comprehension from proactive scaffolding
Delayed independent construction and transfer to unfamiliar tasks; creativity and artifact-quality findings remain mixed or modest
Near transfer: promisingDelayed transfer: unknown
Turning critique, alternatives, and unfamiliar implementation into useful edits
A general “sweet spot”; the direct expertise-stratified sample is too small
PromisingNot settled
Bounded multiplication: option filtering, representation bridging, debugging, and constrained implementation
Open-ended judgment, production delivery, and a universal advantage over direct work
Bounded: moderateField: low
Translation of domain intent into a query, table, code candidate, or familiar chart
Audit of joins, measures, uncertainty, interaction, and generated implementation
Direct evidence: thin
Access is successful new work. Productivity is less total effort to an accepted artifact. Quality uses a declared correctness or usefulness rubric. Learning survives an unassisted transfer test. Verification detects and repairs defects. Reader outcome changes comprehension or decisions. Satisfaction and first-render speed do not stand in for the other rows.
The design requirement is not to pick one “user level.” Expose consequential choices for people with less construction skill, preserve precise control for experienced practitioners, and let every person verify through a representation they understand. One current experiment demonstrates immediate post-removal comprehension from proactive scaffolding; durable construction gain and AI-caused atrophy remain unestablished.
Visualization literacy review ↗ · Who counts as a novice? ↗ · Visualization learning ↗ · Creativity and time ↗ · Current novice study ↗ · Expert replication ↗
Reader experience
Creator delight is not a proxy for reader comprehension. The visual artifact inherits expectations from journalism, science, business, education, or social media.
Compression into marks, labels, annotations, interaction, prose, and provenance
→In a 2025 study, 31 of 37 participants mentioned clarity while ranking charts. Sources, integrity, familiar forms, and aesthetics also mattered, with substantial differences between people. The study measured deliberative trust, not correctness, comprehension, or behavior.
Trustworthy by Design ↗In a preregistered 2021 experiment, people brought strong preferences for human or algorithmic recommendations, but data relevance usually dominated actual selections. The same label suggested precision to some and missing human judgment to others.
Vis Ex Machina ↗On 60 synthetic charts, multimodal models enumerated data structure while 24 people formed trend narratives and reacted more to layout and overlap. Each could succeed on a different notion of intent.
How Do LLMs See Charts? ↗In a randomized 117-person experiment, data stories, passive AI Q&A, and proactive scaffolded dialogue all improved visual-comprehension scores. After support was removed, the proactive group’s median was 6/6 versus 5/6 for both alternatives; completion time did not differ.
GenAI agents and visual comprehension ↗In a 48-person controlled experiment, second-phase accuracy was 71.9% for readers shown AI-generated misleading charts and 88.3% for readers shown correct charts. This measures bounded chart-QA harm—not prevalence, calibrated trust, or consequential decisions.
ChartAttack ↗A 26-person low-vision smartphone study found 100% completion with a full interactive treatment versus 61.5% with a screen magnifier. A separate 10-person blind-reader study found similar device-level accuracy but different time, workload, and chart-type results. Neither tested AI-generated charts.
Context
Preserve the question, data, decisions, artifact, provenance, and acceptance evidence. Then match the authoring and evaluation method to the human purpose.
Helpful: exploration, annotation alternatives, accessible implementation, source checks.
Danger: fluent generation substituting for reporting, authorial judgment, or reader testing.
Helpful: governed measures, stable layout edits, anomaly explanation tied to source visuals.
Danger: silent changes to metrics, thresholds, state, or hierarchy.
Helpful: semantic grounding, visible intermediate data, reusable verification.
Danger: confident answers over weak metadata, ambiguous measures, or hidden filters.
Helpful: many cheap views, branches, undo, direct manipulation.
Danger: turning the first plausible pattern into the final narrative.
Helpful: restricted representations, executable transformations, linked views, domain checks.
Danger: visual plausibility standing in for scientific correctness.
Helpful: adjustable assumptions, guided interaction, feedback, multiple representations.
Danger: generating interactivity without measuring what learners understand.
The lifecycle test
Question → data authority → generation → checking → correction → delivery → reader use → sustainment. The spine stays fixed; acceptance changes with the environment.
Current diaries expose source work, decomposition, defects, correction, mobile checks, publishing, and proposed updates.
Independent verification, editorial review, reader comprehension, accessibility, later refreshes, and civic use.
Source-to-claim trace, editorial gate, delivered desktop/mobile/access states, audience test, correction policy, and update owner.
Product contracts expose semantic models, queries, permissions, and review surfaces; testimony explains the need for narrow, owned metrics.
Routine task success, review queues, refusal, one-number incidents, trust recovery, and total cost in one independent study.
Approved metric contract, exact query and filters, permissions, owner, refusal behavior, escalation, and decision outcome.
One 117-person experiment measures immediate comprehension; one classroom case shows direct instruction rescuing failed AI-assisted construction.
Delayed construction, unfamiliar transfer, ordinary assistant use, and a complete learner-to-reader project.
Immediate and delayed unassisted performance, diagnosis, transfer, confidence calibration, and the eventual audience’s comprehension.
Controlled adjacent studies measure low-vision smartphone and blind nonvisual outcomes; other sources expose generated-description failure and newsroom scale.
AI-assisted authoring joined to disabled readers on their own devices, with verification, recovery, and harm measured.
Co-designed alternatives, real assistive technology, success, error, time, workload, independent verification, feedback, and remediation.
A practical evaluation
A prettier first render is not enough evidence for a workflow change.
Exploration, governed BI, operations, public explanation, science, education, or a reusable application carry different obligations.
Include self-owned data, awkward semantics, a correction request, and the actual delivery surface. Retain a simple task as a control.
Measure human time, errors, correction path, and delivery effort before claiming a gain.
Count prompting, waiting, cleanup, verification, publishing, and monetary cost.
Make fields, filters, aggregation, transformations, code or semantic query, and generated interaction state visible.
Recompute values, exercise controls, check exports, accessibility, responsive behavior, and future maintenance.
Ask defined readers for the claim, evidence, uncertainty, and next action. Measure correctness, time, confidence, and harmful misreadings.
The updated gap ledger
“Partly answered” means one bounded study exists—not that the field can generalize.
How to read the evidence
Supports claims about its participants, tasks, models, measures, and comparisons. Small or dated studies do not automatically transfer to current production.
Shows which jobs and tools respondents report. Self-selection, missing items, and sponsor or community channels limit population claims.
Shows what the provider says the feature can see and do. It does not independently establish correctness, usability, or adoption.
Surfaces jobs and failure costs that experiments may omit. Self-selection means it cannot establish prevalence or comparative performance.
Evidence cut 14 August 2026. Twenty-one human studies and structured workplace evaluations were read in full alongside practitioner surveys, current provider and open-source documentation, public discussions, first-person production cases, and dated essays and interviews from several professional positions. The strongest counter-reading is that better models will erase older failures. Fresh evidence shows real improvement in visual output—and new failure surfaces from richer, slower, more complex artifacts. Capability is moving; the location of human work is moving with it.
Publication history
Initial public edition mapping practitioner jobs, tool choices, creator experience, delivery failures, reader reactions, and open evidence gaps.