Research source document · Evidence reviewed through August 18, 2026

Executive summary: AI-assisted data visualization in August 2026

Status: state-of-the-field snapshot. Evidence cut: 2026-08-16. Recheck by: 2026-11-16, or after a material model, harness, benchmark, registry, or analytics-assistant release.

This summary draws on the research synthesis, the practitioner and reader report, the human-skills and banked-gains review, the learning-with-AI chapter, the agent-skill deep dive, and the specialized-vision deep dive. The capability forecast states which gains would unlock the remaining levels, separates model and technique effects, and records seven dated resolution tests through August 2029. The editorial architecture shows how these can become one comprehensive report and a set of self-contained essays without duplicating the evidence base.

Visual entry points: explore the research field guide, the current and historical capability scorecards, the practitioner and reader experience, or the map of the complete research package.

The short version

AI can now produce useful charts, chart code, scientific views, and increasingly ambitious interactive dashboards. It can accelerate first drafts, translate between unfamiliar representations, automate repetitive implementation, and make visualization accessible to people who could not otherwise build one.

That is not the same as reliably producing an accepted visualization. The remaining failures are often the ones that matter most: hidden transformations, wrong business meaning, regressive corrections, interaction state, mobile and assistive delivery, maintenance, and whether a defined reader understands the claim. A fluent chart can hide these failures rather than expose them.

The field is not converging on one agent, product, or universal workflow. It is converging on a set of controls:

  1. Declare context: purpose, audience, data contract, delivery surface, and stakes;
  2. Expose the work: transformations, encodings, state, and revision history;
  3. Compute what can be computed: use deterministic methods for claims that can be checked deterministically;
  4. Test the delivered artifact: inspect and exercise the actual output; and
  5. Keep consequential acceptance human: reserve contextual, consequential, and reader-level acceptance for people with the relevant authority or experience.

The controls now have a bounded system-level denominator. Six held primary authoring or analysis systems all separate generation from execution and expose some critique route; four expose a meaningful human control gate. Zero of six joins an exact configuration to accepted delivery, intended-reader outcome, or later maintenance. Because their critics range from typed validators and runtime errors to visual models and human inspection, the reference anatomy is an eleven-stage information-and-authority checklist—not a claim that one agent topology wins.

Public acquisition is not adoption. Twelve named data-visualization skill listings carry 37,720 listing-summed install signals at the held evidence cut, but collapse to ten parent repositories and eleven documented provenance lineages. The official CLI’s install event does not observe whether an agent later invoked the skill or whether work improved. Across the twelve named public rows, real-task invocation, repeat use, organizational acceptance, and outcome evidence are all 0/12 observed. Private use remains unknown. Use public counts to choose what to inspect—not to justify procurement, effectiveness, or field-adoption claims.

Available controls are not governed deployment. Power BI/Fabric, Tableau, and Looker expose meaningful enablement, data-boundary, and monitoring controls; two provider stories separately report named-feature organizational use, and two longitudinal studies add adjacent governance-process evidence. Across those seven held rows, zero joins one exact feature to authorization, effective configuration, audit, routine use, incident disposition, an independently accepted outcome, and later recheck. Treat controls and use reports as two rails that must meet on one nine-receipt deployment record—not as a governance score assembled across vendors or customers.

A human study is not automatically delivered-reader evidence. Across ten held primary rows, ChartAttack alone directly carries selected AI-generated charts into an independent controlled-reader effect. Five other rows measure adjacent trust, comprehension, decision, or access outcomes where AI labels, explains, assists, or sits beside the visualization; three concern creators or co-designers; one uses a model reader only. Zero binds an immutable artifact and model to accepted delivery, representative intended-context outcome, and a later same-lineage recheck. Name the AI role before translating the outcome.

Longitudinal evaluation is a development receipt, not audience impact. The 2026 Lexara study follows six CVA developers using a deployed evaluation toolkit for two weeks across 38 experiments, 57 newly authored cases, ten models, and six prompts. They recorded model/prompt choices, rationale, outputs, observations, and confidence; concrete findings changed or challenged development selections. The study does not join an immutable build and exact configuration to an accepted downstream visualization, intended-reader decision, calibrated trust, whole cost, and later post-release recheck. Adding the same-feature InfoViP operational record to the ten reader cases, three production afterlives, and Lexara expands the audit to 15 named cases; zero clears all twelve receipts. Use evaluation exports to inform a frozen acceptance gate; do not call them proof of deployment impact.

A production component is not a complete release afterlife. The same InfoViP line now runs from seven FDA Safety Evaluators’ requirements and prototype feedback to a six-month internal real-work evaluation: approximately 60 reviewers were invited, 20 unique reviewers made 58 submissions, and 27 feedback files returned. The final 2025 CIOMS report says the deduplication component was approved for historical and live ETL, installed in AWS, integrated with the agency AERS, and by 30 July had screened more than 30 million historical reports while continuing to process about 8,000 submissions daily. It says case-series use is fully implemented and administrators continuously monitor the pipeline and its output. It also says a solid QA plan was not yet in place, periodic audits were planned, routine-use roles were incomplete, and downstream signal-discovery effects had not been investigated. Three official FDA award rows total $2.40 million obligated and $2.86 million estimated—a bounded external envelope, not whole cost or ROI. The current federal inventory still lists acquisition/development, implementation N/A, and no AI-system ATO. Keep both sides: this is a self-reported approved and production- integrated component with a bounded routine case-series path, not proof of an immutable accepted whole-platform build, completed QA audit, authorization, current user denominator, attributable outcome, affected-reviewer return, or whole cost. A 2026 FDA fellowship record proposes Elsa GenAI API and interface work; it is a change trigger, not an observed release.

Search recency is not an event date, and a platform launch is not an application join. An FDA CDER Conversation that surfaces with a 2026 published-time field explicitly says its content is current to 3 April 2024; its future-tense InfoViP wording therefore precedes rather than contradicts the final 2025 component report. A May 2026 FDA biography still names Oanh Dang as serving as InfoViP project lead, which is not operations or maintenance authority. FDA’s separate Elsa 4.0/HALO release says Elsa launched agency-wide and HALO integration began, but never names InfoViP. Keep the specific integration prospective until a named build, release, exposure, or use receipt appears.

Three research mentors are not an accountable operating chain. The live InfoViP–Elsa opportunity is now marked closed after its 8 May deadline and names Joshua Xu, Leihong Wu, and Oanh Dang as contacts for the nature of the research. Current FDA profiles coordinate Xu to Research-to-Review and Return and advanced-AI integration, Wu to AI/ML bioinformatics research, and Dang to InfoViP project leadership. None assigns InfoViP service operation, maintenance, QA execution, release approval, or authorization. The opportunity also makes a fellow a nonemployee barred from inherently governmental functions. Treat closed as an administrative state—not selection, execution, approval, or release.

Audience-conditioned generation is not audience testing. A bounded screen of the review’s 122 authoritative stable-key titles surfaces one explicit GPT study: Beyond Generating Code evaluates 91 quiz questions and nine homework assignments, and its one-commit supplement pins prompts, outputs, code, and grader bundles. That is a real custody upgrade. It is not an immutable provider snapshot, a versioned release, or evidence that the government, UN, or primary-school audiences named in prompts received or evaluated the artifact. There is no later commit or audience recheck. Report this as one pinned-output bundle and zero delivery/recheck chains. The other 121 title nonmatches are not exclusions, and the complete denominator remains unknown.

A mature visualization afterlife does not validate an AI feature. Exact- DOI discovery now covers 114 authoritative review keys and 87 available abstracts. It adds no hidden GenAI signal, but it surfaces AIDSVu: ten years of public delivery, broad aggregate use, named planning applications, governance, and a later 2026 data release. Those are strong platform receipts. The held surfaces expose no AI role, immutable build, version-bound measured audience outcome, or affected-audience recheck. Keep AIDSVu as a non-AI comparator. Do not borrow its reach and afterlife into a model claim, and do not turn 86 abstract cue nonmatches into exclusions.

A human study is not deployment. Primary-text recovery below the 27 no-abstract DOI rows puts five exact texts into custody: four empirical papers and one review. The empirical papers separate controlled audience measurement, public-sector co-design/demo/bounded use, and enterprise demonstrator plus production intent. None binds an accepted field release to a later affected- audience recheck, and all five explicit GenAI cue screens are empty. Keep 22 uncaptured DOI rows and eight non-DOI keys unassessed; never join the task, demo, and production-intent receipts across artifacts.

A field-use receipt is still not an AI lifecycle. The eight stable keys below A208’s DOI layer were partly an authority problem: exact publisher records recover three DOIs, correct two publication years, and show that one printed PMID names a different paper. Four full texts, two primary abstracts, one issue excerpt, and one gated primary text are now distinguished. The strongest practice receipt is a 2012 EMR cancer diary: its abstract reports increasing use in system logs, while an independent systematic review recovers an 11-clinician QUIS survey with median satisfaction of 4.38. That is meaningful embedded-use and usability evidence. It exposes no AI role, immutable accepted build, patient or decision outcome, later material event, or affected-clinician recheck. Keep the case as a non-AI operational comparator, keep the gated row unassessed, and keep the complete-review lifecycle denominator unknown.

Seven human evaluations are still not one deployed lifecycle. A targeted recovery of nine high-signal DOI rows adds eight substantive primary abstract or official-project surfaces and seven explicit participant/evaluator studies, leaving 14 DOI rows without substantive primary content. The strongest delivery near-miss is InfoViP: seven FDA safety evaluators shaped and evaluated the prototype, suggestions were addressed, and an official FDA page says an enhanced NLP and unsupervised-learning version will be installed in production. That is prospective. No held receipt confirms installation, accepted release, routine use, regulatory outcome, later change, or evaluator return. Keep the case as a non-generative production-intent comparator.

The remaining DOI content gap is one record, not a deployment finding. All 14 residual rows were investigated; eleven gained substantive primary content in the residual pass, including one full publisher chapter, while B92, B91, and B110 were then metadata-only. The chapter connects three UX experts and 25 problems to a third design version. A corridor stakeholder case and university-network case add practice context. None has an explicit GenAI visualization role, accepted field release, routine-use denominator, consequential outcome, or later affected-audience recheck.

One named interface is live; that is delivery, not observed use. A final- three recheck gives B110 an official abstract, pinned framework source, and a DiscoverWater interface reachable at KU in August 2026. At that pass, B92 and B91 remained content-unassessed. No manifest binds the live bytes to a commit, and no held receipt shows acceptance, continuous or ordinary use, consequential audience outcome, later affected-audience return, whole cost, or an AI visualization role.

Public source only partly reproduces what is live. The sole published v1.2 page state does not match the dated live page, and the linked v2.0 source is an R/Shiny application. Eight of twelve live dependency names occur in pinned application source and two of three sampled asset pairs match canonically, so partial lineage is real. It still supplies zero exact-build or measured audience-outcome receipts; a prototype demonstration and analytics hook do not change that state.

One exact publisher abstract closes one content gap, not a lifecycle. B92’s publisher record describes a multicriteria hydrogen-pipeline risk model, Monte Carlo simulation, Kendall’s tau rank comparison, and graphs for ranking sections and targeting mitigation. The abstract names no human sample or AI visualization role and supplies no release, delivery, observed-use, outcome, later recheck, or whole cost. B91 remains content-unassessed, so the current 27-row ledger is six full texts, 20 primary abstract or official-content rows, one unassessed row, and zero complete AI visualization lifecycles. Three bounded lawful passes have exhausted the same six named B91 routes, so generic retrieval is paused with four exact reopen conditions. That is one access-blocked row and zero active generic B91 targets—not a negative finding or permission to drop it.

A review corpus is not automatically a lifecycle denominator. One 2025 systematic review maps 127 studies of visualization in broadly defined AI-assisted decision-making: 118 empirical papers and nine reviews. It explicitly does not focus on GenAI and does not jointly extract immutable artifact, delivery, consequential decision or calibrated-trust outcome, and later recheck. The complete GenAI lifecycle count is therefore unknown—not zero of 127. Use the review for coverage and primary-source discovery, not as a negative denominator. Its own publisher supplement adds a prior gate: the six domain totals sum to 127, but the “full list” names 126 citations and leaves one Education entry as a literal question mark. The review’s own post-selection Education paragraph and bibliography now reconstruct that row uniquely as Hernández-Calderón et al. (2023), with matching DOI metadata. Keep the published ? beside the repair overlay. The completed six-domain crosswalk now binds 122 of 127 supplement positions to distinct A208 review keys. Five printed rows have exact-looking external DOI candidates but no review-controlled join: Jiao (2022), Burt et al. (2017), Theis et al. (2018), M. Lu (2020), and Islam et al. (2022). They remain candidates, not authoritative keys. One admitted key also carries a printed PMID that resolves to a different paper. The result is 122 review-owned keys, five authority gaps, one quarantined identifier conflict, and an unknown lifecycle count; screening has not started.

These controls are not a validated best configuration. A five-threat audit found no current Vizier configuration result to promote. Common-task studies show why: models can score strongly without reproducing human error patterns or human interpretive strategies. A benchmark therefore licenses only its named task-and-grader lane. Moving to overall quality, audience fit, comprehension, accessibility, accepted delivery, maintenance, or an architecture winner requires a new receipt at that layer. Older human task definitions may remain useful, while the exact model and harness result must be refreshed after material releases. The current qualifying configuration-run count is zero.

Natural language is becoming one control surface inside a structured environment—not a replacement for the environment.

For a governed-BI pilot, “declare context” now has a concrete minimum: name what the assistant can traverse beyond the visible report, the effective identity and permission state, the semantic objects included or excluded, the expected answer or refusal, and the owner of reproduction, correction, and regression. One captured practitioner issue and one Microsoft-authored customer account make those questions concrete, but remain C-grade self-report. They do not establish a product bug, security breach, general efficacy, prevalence, or demand. See the audience decision report.

Capability improved quickly, but the target also became harder

From 2023 through August 2026, general models gained image understanding, code execution, tool use, browser operation, persistent workspaces, and reusable skills. Research advanced from single-chart generation and chart-to-code tasks to editable multi-view dashboards, replayable interaction, realistic chart reading, visual integrity, and expert-adjudicated critique.

Within individual benchmarks, later systems often make large gains in visual fidelity and task breadth. Execution and semantic reliability improve less uniformly. Across benchmarks, there is no honest single progress curve because the inputs, models, renderers, judges, and definitions of success change. Stronger baselines also expose harder questions: not merely whether the chart renders, but whether the data binding is right, the interaction works, a repair introduces a regression, or a reader learns the intended point.

Two August benchmarks sharpen that boundary. Dashboard2Code now tests stateful reconstruction across 180 Plotly Dash applications and 450 interaction tasks; hidden state and transformations can be wrong while the interface responds. Chartography’s best tested configuration reaches 45% mean pass@1 on 100 deliberately difficult, practitioner-authored professional chart-reading tasks. The latter is not an ordinary-chart failure rate. Together they show that interactive state and professional visual conventions are distinct capability frontiers, not details covered by a generic screenshot score.

DV-World adds the next pre-delivery rung without proving production maintenance. Across 50 native Excel chart repairs, the best reported agent route succeeded 48% of the time; across 80 changed-data or evolving-requirement tasks, the top overall score was 51.44%. Those tasks make binding failures and destructive regressions testable. They still end at one benchmark-scored candidate: no deployed artifact, elapsed later event, maintainer, reader outcome, direct-work baseline, or whole cost was observed.

Three same-artifact production afterlives now exist, with different evidence strengths. A four-month KubeStellar Console experience report, commit-pinned quality record, and linked pull requests follow one AI-assisted dashboard codebase through a live runtime regression, repair, a separate build blocker, restored deployment, and a manually added prevention check. That is a real rung beyond DV-World’s prepared changes. It is one solo-maintainer project-authored case, not an independent productivity comparison or evidence of safe autonomy.

OpenClaw Agent Dashboard adds a second crossing and the first here with a non-owner operator report. The project says it was built with Claude Code and released v3.0.0 on March 5, 2026. Thirteen days later, an operator reported a cost view showing $0 for a custom provider despite present token data; the owner acknowledged that the configuration had not been tested. The captured issue remains open. This is distribution-to-failure evidence. A different non-owner has since opened PR 31, which proposes a provider-inference fallback matching the reported mechanism. The proposal has no maintainer review, merge, corrected release, or reporter recheck, so it is not recovery. Project-level AI attribution does not identify who wrote the defective line.

Prism adds the first held multi-attempt delivery-and-repair chain. Its maintainer says the family dashboard was built entirely with Claude Code under human product direction. In issue 81, a non-owner operator reported an initial Home Assistant add-on install failure, a failed repair build, and then install success followed by startup failure; the operator used ordinary Docker instead. The later v1.8.11 release and release-pinned entrypoint repair the recorded socket and schema-replay failures, and the maintainer reports install, start, and restart checks on real Home Assistant OS. The affected operator never rechecked that release. This is repaired-release and maintainer-verification evidence, not independent recovery, and it does not attribute any defect to Claude Code.

The pull-request histories add a separate, bounded form of handoff evidence. Prism’s non-owner PR 23 carried a Claude-assisted weather visualization through maintainer-found delivery defects, contributor repairs, owner integration, and a credited release; the contributor returned with accepted PR 41. OpenClaw’s non-owner PR 15 put Claude-assisted multi-provider pricing into the product after maintainer integration. These are receipts for accepted second-person change, and Prism adds repeat contribution. They do not name a second maintainer, transfer release or incident authority, or independently accept either later recovery.

The consolidated production result is short: three afterlives, two maintainer-verified restorations, zero affected-user or independent-operator recoveries, zero transferred maintenance authority, and zero complete twelve- state cases. A closed issue is not a recovery receipt, a maintainer check is not an affected-user check, and an accepted contribution is not operational authority. The next evidence worth funding is a return-user exercise of the corrected route or a named second person actually releasing or responding to an incident.

A second single-case PM4Py-UCM experience report measures about 65 active hours across 18 agent sessions and ten weeks; fixes outnumbered features 2.3 to 1, and 78% of visualization and layout work was fixes. It makes refinement cost visible without following a later field incident or reader. Neither case supplies the complete human-and-model cost, maintenance-authority handoff, accessible-reader use, consequential decisions, or calibrated trust needed to close the lifecycle question.

A project history is not a same-artifact afterlife. A closed audit of four promising public candidates admitted zero new episodes. The strongest near-miss is MAIDR: 11 blind participants tested four chart types using their own screen-reader and refreshable-braille-display paths, but the studied implementation predates AI. The current project declares a complete TypeScript rewrite and an AI description layer; no captured source pins a representative participant recheck to that release. Separately, a public-health copilot paper describes a 16-person trust/usability method but says the experiment was removed and reports no human result. Methods, results, release, and later use remain separate gates.

The public legacy code now narrows the later AI study further without closing that gate. A dedicated box-plot study surface is present before v2.10.0, making that tag the first containing version floor found—not proof of the build or model used by the eight blind and low-vision participants named in the abstract. Nine selected later model, verification, and deprecation events add real maintenance history but no representative-user recheck. The separate TypeScript rewrite needs its own validation.

The executive decision rule is: preserve the exact artifact from acceptance through deploy, incident, repair, corrected release, independent recheck, restored service, and the new regression gate—then count the human infrastructure and follow the result to a second maintainer and intended readers. Track accepted change, repeated contribution, maintenance authority, and independent recovery as separate states. Treat unknown, zero, and not applicable as different states.

The benchmark crosswalk shows exactly what each major evaluation measures and cannot establish. The capability scorecard turns that evidence into current, historical, and forecast scores by use context.

The forecast derived from this record is deliberately asymmetric. It assigns higher near-term probabilities to accepted static work, interactive-dashboard reasoning, narrow visual critics, and environment-specific adapters because they have public tests and observable failure signals. It assigns lower probabilities to production productivity, representative-reader benefit, and autonomous publication because the evidence needed to verify those outcomes is mostly absent. Better models can help request and reason over local definitions, provenance, authority, and reader response; they cannot manufacture those facts.

The first forecast check resolves nothing—the earliest date is August 2027—but it does correct the record. ViviDoc was published before the forecast freeze and captured afterward, so it is disclosed as a pre-freeze source omission rather than a new capability event. Its typed State–Render–Transition–Constraint plan improved same-pipeline interactive-document measures across three backbones. It did not measure representative readers, accessible delivery, whole production cost, later maintenance, or autonomous publication, so all seven forecasts remain open and unchanged.

What consistently helps

These are shared principles, not a single architecture. Journalism, governed BI, operational monitoring, exploration, science, education, and reusable applications have different purposes, authorities, update cycles, and failure costs. A common record of intent, data, transformations, artifact, revisions, and evidence should route to different authoring and acceptance procedures.

The specialist-cost question is now measurable, but not measured. Nine primary studies report fragments: roughly 4–16× inference for repeated chart extraction; about 31–33 seconds of human verification per chart in ExChart; about 6–10 seconds for one GPT-4o call versus 90 seconds serial or 30 seconds parallel and about $0.40 per query for ChartAgent; substantial crowd and expert work in VisJudge-Bench; and 450 manually annotated dashboard interactions plus human validation in DashboardMimic. A sixth chart-specific routing paper separates training resources from per-chart inference, but its displayed split totals, latency percentage, memory-delta field, and selected configuration do not reconcile; its exact numbers are quarantined. METAL adds same-task quality results, but its five-candidate baseline makes five generation calls while its five-recurrence multi-agent route can make one initial generation plus fifteen critique or revision calls. Neither its paper nor released logs preserve the per-arm usage needed to reconcile that difference. A new ChartAgent tool-integrated-reasoning preprint reports same-task accuracy beside 3.0–9.8 mean tool invocations across scheduler settings. The count omits model and local-tool compute, reflection, tokens, latency, retries, evaluator and human work, and task-level accepted outcomes. None of those eight visualization fragments joins one current-general and specialist route through acceptance, delivery, readers, and maintenance, so they cannot be added into a route total.

The ninth fragment is adjacent rather than visualization-specific. ProMCP shows measured token and latency bottlenecks moving between planning/schema injection and final synthesis as MCP client topology changes. Its pinned release has useful event fields but no task traces or historical run manifest, and its benchmark entry point is not runnable as released. The result adds cold start, discovery, phase, topology, cache, streaming, concurrency, retry, and observability state to the receipt; it does not add a visualization baseline or cost winner.

The decision rule is an eleven-lane trial: start with route preparation and ownership, then keep task/acceptance, success, inference, latency, evaluation, human correction, unresolved failures, maintenance, delivery, and reader outcomes separate under one equal total budget. Freeze a multidimensional cap, then retain observed calls and retries, input/output tokens or image units, evaluator and tool work, accelerator use, latency and concurrency, charges, and human minutes for every task. Matching rounds, candidates, rates, nominal caps, or model labels is not an equal-budget receipt. Distinguish fixed preparation, periodic ownership, marginal attempts, failure-contingent work, and downstream delivery or reader cost; declare the volume and time horizon before amortizing. Within machine work, retain initialization and discovery separately from planning, tool calls and results, context updates, and answer synthesis; bind each event to the effective host/client/server/model/tool topology and runtime state. Hidden internal phases stay missing rather than inferred. Recompute source denominators and percentages before admitting telemetry, then report cost per eligible, attempted, accepted, and reader-successful task. This turns the gap into an executable procurement and research question without claiming that a specialist or general route is cheaper.

Cost per accepted artifact now has an explicit two-receipt rule. Use all reconciled route cost over the declared population and observation window as the numerator, including rejected, abandoned, quarantined, and no-output work. Use only artifacts passing the frozen human or production acceptance contract as the denominator. Tool calls, per-attempt API averages, aggregate accuracy or F1, finish signals, and one-arm telemetry are not substitutes. Zero of the twelve held fragments supplies both receipts.

One result now clears a narrower bar. The peer-reviewed Selective TTS visual- insights study matches a declared LLM-call budget and, in a separate experiment, output tokens inside one pipeline. Its released code still omits repair and errored-worker cost, records completion tokens only, and counts proxy-judge-scored candidates rather than contract-accepted artifacts. Treat this as a matched declared partial inference budget—not observed total use or cost per accepted work.

A second result makes retry burden visible without pricing it. On 888 VisPlotBench tasks, VisCoder2-32B begins with 649 execution passes and GPT-4.1 with 563; after up to three conditional self-debug rounds, both finish at 732, or 82.4%. The paper tables reconstruct to 584 versus 714 revisions, with diminishing rescue in later rounds. The released debug paths do not retain comparable task-level tokens, compute, latency, or charges, and execution pass is not human or production acceptance. Read this as same endpoint, different retry topology, no cost winner—and persist resource telemetry before using a retry count in procurement.

What creators and readers actually experience

The practitioner and reader experience explains why enthusiasm and frustration coexist. First drafts can appear before an idea has cooled, and tedious implementation can compress dramatically. Correction and verification can erase that gain. In one study of 18 experienced analysts completing 108 deliberately error-prone tasks, seven episodes were unfinished and participants prematurely accepted the result 31 times. A novice study found low verification intent and repeated failed repairs alongside positive satisfaction.

Expertise changes both the benefit and the burden. Experienced practitioners can filter weak suggestions and use AI selectively, but they also notice when describing a small edit takes longer than making it. In interviews with 17 biomedical-visualization practitioners, AI was used chiefly for auxiliary work; all opposed substantial generated visuals in final scientific communication.

Expertise is better represented as a capability profile than a rank. Established visualization-literacy research distinguishes consuming, constructing, critiquing, and connecting a chart to its context; professional work adds data, domain, implementation, situated-judgment, and delivery resources. Current learner studies bank access, reported speed, confidence, and engagement more readily than correctness, verification, or delayed independent skill. One 117-person randomized visual-comprehension study found an immediate post-removal advantage for a proactive agent that asked scaffolded questions; it did not test delayed construction or far transfer. Adjacent randomized education studies show why the distinction matters: assisted exercise performance can improve while conceptual learning does not, and unrestricted answer access can harm later unassisted work. Small studies suggest intermediate and expert practitioners can turn structured critique and constrained implementation into better work more reliably, but no large study establishes one universal expertise curve.

The decisive learning study is now concrete, but still unrun. It would compare conventional instruction, answer-oriented AI, and metacognitive AI using the same frozen model and harness; remove assistance immediately, at six weeks, and at six months on unfamiliar tasks; then test the resulting artifacts with blinded readers on their own devices and assistive paths. Ten outcome lanes keep access, assisted performance, learning, cost, accessibility, and reader use separate. No owner, ethics approval, powered sample, participant, budget, or result exists, so this design narrows the research gap without answering it.

Readers receive the claim, not the authoring transcript. In a controlled 48-person study, selected AI-generated misleading charts reduced answer accuracy from 88.3% to 71.9%. Adjacent accessibility studies show why delivery cannot be inferred from a responsive screenshot or text alternative: outcomes for low-vision and blind readers changed with the actual device, interaction, modality, chart type, time, and workload. A new 12-person blind and low-vision study adds direct AI-assisted learning evidence: eleven preferred tactile charts plus text and chat and described a stronger spatial model, but measured chart-understanding accuracy did not improve. No current study joins AI generation to a representative mobile or assistive-technology reader evaluation.

Following the artifact past its creator exposes another class of failure. In one public second-maintainer episode, an AI-built dashboard failed during the creator’s absence because refresh and scheduling lived on that person’s laptop; the inheritor replaced it with a proper pipeline. This is testimony, not a prevalence estimate, but it makes execution host, lineage, definitions, dependencies, documentation, ownership, and rollback part of acceptance—not optional handoff cleanup.

Skills can help; popularity does not say whether they do

The agent-skill deep dive finds that the public ecosystem contains hundreds of listings labeled as data visualization, with extensive copying, vendoring, and source drift. Registry installs measure acquisition. Repository stars usually belong to a much larger repository. Neither measures routine use or output quality.

The inspected packages perform six different jobs: primers, storytelling guides, renderer gateways, library or domain adapters, quality gates, and structured workflows. Their most defensible value is information a capable model cannot reliably infer: a current API, local environment, semantic representation, known failure, executable validator, or delivered-surface repair loop.

There is direct evidence that this can work. SciVisAgentSkills improved quality in all ten paired suite-by-agent comparisons across 108 scientific- visualization tasks, although one completion measure fell. Broader benchmarks show the boundary: compact, relevant, compatible skills can help; comprehensive, self-generated, stale, or incorrectly retrieved guidance can perform worse than no skill. Popular generic, dashboard, accessibility, mobile, and explanatory packages still lack independent paired evaluation.

Specialized vision is a sensor, not a final judge

The specialized-vision deep dive finds that “vision for charts” includes at least six jobs: question answering, structured extraction, OCR and layout, element grounding, integrity or perceptual critique, and verification or repair. No model covers all six reliably.

Specialists show real component gains. VisJudge-7B matched its expert- adjudicated quality rubric better than the tested general models. ChartAgent’s chart-specific tools materially improved the same general reasoner on numeric questions. Chart-aware grounding and recent OCR/layout models improve bounded perception. At the same time, newer out-of-distribution extraction and realistic chart-QA tests show strong general models overtaking older chart specialists. Dashboard2Code adds a stateful fixed-desktop test but leaves responsive/mobile, animation, keyboard, and assistive-technology behavior open. Chartography shows that domain conventions and hard professional visual forms still defeat strong general systems even with more reasoning.

The practical choice is role-based: use source data and deterministic checks whenever available; add a specialist for a measured parsing, localization, or perceptual failure; retain a current general model for broad reasoning; and do not confuse any model’s score with reader comprehension or publication acceptance.

What performs poorly, adds risk, or remains unsupported

The next evidence should follow work to acceptance

The capability forecast and research-package map show where evidence is still missing. The largest gaps sit between measured capability and lived outcome:

A ten-case bridge audit narrows that statement without reversing it. Three controlled studies connect technical evidence to analyst use, analyst correction, or reader harm. Zero audited case binds an exact system version and frozen acceptance rule to real delivery, and zero follows one artifact through reader or decision use, a later maintenance event, and whole cost. The bridge is short, not empty—and it is not a complete lifecycle.

One current readiness audit makes the experimental gate concrete. It reconciled 14 historical case definitions, 18 run directories, and 13 judge directories, and now has all five required context slots with executable source packets. The exploratory packet is a commit-pinned, CC0 Palmer Penguins exploratory fixture: across 342 complete rows, bill length and depth correlate -0.235 when pooled but positively within every species. That reversal can test source, grouping, branching, and render receipts. Because the dataset is familiar and public, it is not held-out efficacy evidence or a human outcome. The third packet is a first-party synthetic governed-BI contract: an owned measure, vintage, four roles, six verified queries, four refusal cases, and an exact fanout incident. A bad tag join returns $16,830 against the authoritative $7,730. That makes semantics, permission, refusal, and recovery testable—not production security or reliability. The fourth packet uses official 2024 ACS B19013 estimates and 90% margins of error. Iowa, Kansas, Montana, and Wyoming differ by only $13 to $192; no internal pair clears the declared 90% comparison threshold. That makes exact uncertainty propagation, refusal of a false winner, and expert review routing testable—not a scientific result across domains or an expert or reader outcome. The fifth packet is a first-party synthetic interactive maintenance explorer with two data versions, one canonical state, ten expected states, seven isolated trajectories, keyboard and narrow-delivery contracts, and a stale-selection failure. When a filter makes T007 ineligible and leaves only T012 visible, detail, accessible summary, URL, and export must clear T007 together. That makes synchronization, replay, recovery, and prepared change testable—not a browser run, accessibility result, human outcome, or production maintenance result. Historical auditability and five-of-five source-packet readiness are not a current baseline: compare no architecture until the exact current model/harness manifest is frozen and every arm produces equivalent receipts.

A first current-harness preflight made that distinction executable. Its seven-blocker predecessor locked the five source packets, local toolchain and available-skill observations, task order, three arm roles, and an 18-field common receipt shape. A credential-free neutral runner now clears one of those blockers: it recomputes five source checks, stages 15 isolated fixture-arm attempts, separates model input from evaluator-only truth, hashes artifact, render, trace, and check files outside the arm, and rejects missing or mismatched receipts. This matters because the legacy editorial case embeds its answer key; handing over the whole packet would leak the evaluator into the task.

The successor preflight still reports six pre-run blockers, zero model calls, and no spend authority: the exact served-model snapshot and server harness are not exportable, no same-model bare interface is available, no one Vizier mechanism or acceptance authority is selected, and no capped model budget is authorized. The passing runner self-test is evaluation plumbing—not a model, artifact, browser, human, or architecture result. A model label and local CLI version remain useful inventory, not a reproducible comparison.

The proposed Vizier delta is now narrow enough to decide without calling it selected. The two public skills each combine framing, retrieval, guidance, tools, inspection, and critique; either whole package would be another bundle comparison. The prepared candidate adds one artifact-evidence reconciliation checkpoint after the first executable artifact and before repair or finalization. It asks the arm to bind visible claims and states to available source, transform, query, calculation, branch, or canonical-state evidence and freeze mismatches before repair. Its corrected schema refuses an overall pass when any surface is mismatched or unresolved, required states are missing, or repair remains necessary. It adds no model, agent, tool, answer key, or call. The Vizier owner has not ratified it, so the layer remains null and the preflight still has six blockers. A testable protocol is not an intervention choice or an effect result.

The acceptance gate is also narrow enough to decide without pretending a model or checklist is a person. One candidate separates deterministic checks from six nonmechanical lanes: context review, domain expertise, accessibility conformance, evaluation with relevant disabled users, intended-user acceptance, and later maintenance handoff. Across five fixtures, a mechanical-only pilot would retain 30 explicit not-run records. A human pass requires a named eligible human authority and evidence; an automated tool cannot occupy that role. W3C’s evaluation overview says no tool alone determines accessibility, while its user-evaluation guidance keeps user experience distinct from conformance and warns against generalizing beyond the participant scope. No acceptance scope is selected and no person was contacted, so the research-owner gate remains open and the blocker count stays at six.

Those six blockers now have one attributable response contract rather than one blanket “go” decision. The register exposes 17 admissible paths across model- provider/runner, harness-provider/runner, research-owner, and Vizier-owner authority. It rejects wrong-role, incomplete, secret-bearing, partial, and six-answer-without-release states. Provider facts cannot authorize spend; credential availability cannot amend a protocol; a passing candidate verifier cannot select a layer or acceptance scope. Even six advancing answers require a separate research-owner release of the resolved decision-set digest, named run plan, call cap, and stop condition. All six selections and release remain null, readiness and spend remain false, and calls and contacts remain zero. The research review translates that same custody state for practitioners, BI leaders, researchers, newsrooms, product teams, educators and accessibility specialists, and broad AI readers.

Future receipts also have a scorer that cannot hide missing evidence inside one grade. It emits computable and checkable mechanical tiers plus six named context, domain, accessibility, user, and maintenance lanes. Empty or not-run mechanical evidence remains incomplete; human authority is bound to the exact lane; overall_score and universal_pass are invalid fields. Its verifier created 15 synthetic fixture-arm ledgers and preserved 90 human not-run states with zero calls. Those are plumbing tests, not successful AI artifacts. The research review shows how practitioners, leaders, researchers, newsrooms, product teams, educators and accessibility specialists, and broad AI readers should interpret the same lane-specific record.

The next cheap hypothesis test is also preregistered but unrun. R4 now has five opaque, hashed visualization inputs: a 19-file model allowlist keeps five visuals, five briefs, and nine source-evidence files separate from five hidden defect records. Two hashed prompts and two fail-closed schemas keep integrity and readability separate; they accept findings, empty completions, and abstentions while twelve adverse records reject aggregate, mixed, identity, evidence, and state failures. It would compare three exact models with one independent human data editor and a predeclared tie rule. Model slots, editor calibration, budget, release, and results remain absent. Five cases are a protocol pilot, not confirmation; no ranking or divergence exists.

The five residual gate categories now have a seven-receipt decision surface: three exact served-model receipts, editor acceptance, same-editor calibration, a capped 30-primary-call budget, and final release over the exact prerequisite digest. Twenty advancing and non-advancing paths validate and sixteen adverse states fail closed. This makes the handoff auditable; it supplies no real provider, editor, spend, release, call, ranking, or result.

The editor-calibration handoff is also concrete but blank: ten baselines, ten synthetic response-quality vignettes with two swapped repeats, three valid response states, and fourteen rejected adverse records. It prevents model output from setting the baseline, but it does not appoint an editor or produce a resolution, tie threshold, score, ranking, or result.

The next experiments should begin with the current model and its normal harness, use representative tasks from declared environments, separate first draft from accepted artifact, add one mechanism at a time, exercise the delivered surface, and retain every correction and regression. Additional instructions, agents, or models should survive only when they improve that full path for a stated job at a proportionate total cost.

Update log

Read or download the Markdown source