Executive summary: AI-assisted data visualization in August 2026
Status: state-of-the-field snapshot. Evidence cut: 2026-08-16. Recheck by: 2026-11-16, or after a material model, harness, benchmark, registry, or analytics-assistant release.
This summary draws on the research synthesis, the practitioner and reader report, the human-skills and banked-gains review, the learning-with-AI chapter, the agent-skill deep dive, and the specialized-vision deep dive. The capability forecast states which gains would unlock the remaining levels, separates model and technique effects, and records seven dated resolution tests through August 2029. The editorial architecture shows how these can become one comprehensive report and a set of self-contained essays without duplicating the evidence base.
Visual entry points: explore the research field guide, the current and historical capability scorecards, the practitioner and reader experience, or the map of the complete research package.
The short version
AI can now produce useful charts, chart code, scientific views, and increasingly ambitious interactive dashboards. It can accelerate first drafts, translate between unfamiliar representations, automate repetitive implementation, and make visualization accessible to people who could not otherwise build one.
That is not the same as reliably producing an accepted visualization. The remaining failures are often the ones that matter most: hidden transformations, wrong business meaning, regressive corrections, interaction state, mobile and assistive delivery, maintenance, and whether a defined reader understands the claim. A fluent chart can hide these failures rather than expose them.
The field is not converging on one agent, product, or universal workflow. It is converging on a set of controls:
- Declare context: purpose, audience, data contract, delivery surface, and stakes;
- Expose the work: transformations, encodings, state, and revision history;
- Compute what can be computed: use deterministic methods for claims that can be checked deterministically;
- Test the delivered artifact: inspect and exercise the actual output; and
- Keep consequential acceptance human: reserve contextual, consequential, and reader-level acceptance for people with the relevant authority or experience.
The controls now have a bounded system-level denominator. Six held primary authoring or analysis systems all separate generation from execution and expose some critique route; four expose a meaningful human control gate. Zero of six joins an exact configuration to accepted delivery, intended-reader outcome, or later maintenance. Because their critics range from typed validators and runtime errors to visual models and human inspection, the reference anatomy is an eleven-stage information-and-authority checklist—not a claim that one agent topology wins.
Public acquisition is not adoption. Twelve named data-visualization skill listings carry 37,720 listing-summed install signals at the held evidence cut, but collapse to ten parent repositories and eleven documented provenance lineages. The official CLI’s install event does not observe whether an agent later invoked the skill or whether work improved. Across the twelve named public rows, real-task invocation, repeat use, organizational acceptance, and outcome evidence are all 0/12 observed. Private use remains unknown. Use public counts to choose what to inspect—not to justify procurement, effectiveness, or field-adoption claims.
Available controls are not governed deployment. Power BI/Fabric, Tableau, and Looker expose meaningful enablement, data-boundary, and monitoring controls; two provider stories separately report named-feature organizational use, and two longitudinal studies add adjacent governance-process evidence. Across those seven held rows, zero joins one exact feature to authorization, effective configuration, audit, routine use, incident disposition, an independently accepted outcome, and later recheck. Treat controls and use reports as two rails that must meet on one nine-receipt deployment record—not as a governance score assembled across vendors or customers.
A human study is not automatically delivered-reader evidence. Across ten held primary rows, ChartAttack alone directly carries selected AI-generated charts into an independent controlled-reader effect. Five other rows measure adjacent trust, comprehension, decision, or access outcomes where AI labels, explains, assists, or sits beside the visualization; three concern creators or co-designers; one uses a model reader only. Zero binds an immutable artifact and model to accepted delivery, representative intended-context outcome, and a later same-lineage recheck. Name the AI role before translating the outcome.
Longitudinal evaluation is a development receipt, not audience impact. The 2026 Lexara study follows six CVA developers using a deployed evaluation toolkit for two weeks across 38 experiments, 57 newly authored cases, ten models, and six prompts. They recorded model/prompt choices, rationale, outputs, observations, and confidence; concrete findings changed or challenged development selections. The study does not join an immutable build and exact configuration to an accepted downstream visualization, intended-reader decision, calibrated trust, whole cost, and later post-release recheck. Adding the same-feature InfoViP operational record to the ten reader cases, three production afterlives, and Lexara expands the audit to 15 named cases; zero clears all twelve receipts. Use evaluation exports to inform a frozen acceptance gate; do not call them proof of deployment impact.
A production component is not a complete release afterlife. The same
InfoViP line now runs from seven FDA Safety
Evaluators’ requirements and prototype feedback to a six-month internal
real-work evaluation: approximately 60 reviewers were invited, 20 unique reviewers made 58
submissions, and 27 feedback files returned. The final 2025 CIOMS
report
says the deduplication component was approved for historical and live ETL,
installed in AWS, integrated with the agency AERS, and by 30 July had screened
more than 30 million historical reports while continuing to process about
8,000 submissions daily. It says case-series use is fully implemented and
administrators continuously monitor the pipeline and its output. It also says
a solid QA plan was not yet in place, periodic audits were planned, routine-use
roles were incomplete, and downstream signal-discovery effects had not been
investigated. Three official FDA award rows total $2.40 million obligated
and $2.86 million estimated—a bounded external envelope, not whole cost or
ROI. The current
federal inventory
still lists acquisition/development, implementation N/A, and no AI-system
ATO. Keep both sides: this is a self-reported approved and production-
integrated component with a bounded routine case-series path, not proof of an
immutable accepted whole-platform build, completed QA audit, authorization,
current user denominator, attributable outcome, affected-reviewer return, or
whole cost. A 2026 FDA fellowship
record
proposes Elsa GenAI API and interface work; it is a change trigger, not an
observed release.
Search recency is not an event date, and a platform launch is not an application join. An FDA CDER Conversation that surfaces with a 2026 published-time field explicitly says its content is current to 3 April 2024; its future-tense InfoViP wording therefore precedes rather than contradicts the final 2025 component report. A May 2026 FDA biography still names Oanh Dang as serving as InfoViP project lead, which is not operations or maintenance authority. FDA’s separate Elsa 4.0/HALO release says Elsa launched agency-wide and HALO integration began, but never names InfoViP. Keep the specific integration prospective until a named build, release, exposure, or use receipt appears.
Three research mentors are not an accountable operating chain. The live InfoViP–Elsa opportunity is now marked closed after its 8 May deadline and names Joshua Xu, Leihong Wu, and Oanh Dang as contacts for the nature of the research. Current FDA profiles coordinate Xu to Research-to-Review and Return and advanced-AI integration, Wu to AI/ML bioinformatics research, and Dang to InfoViP project leadership. None assigns InfoViP service operation, maintenance, QA execution, release approval, or authorization. The opportunity also makes a fellow a nonemployee barred from inherently governmental functions. Treat closed as an administrative state—not selection, execution, approval, or release.
Audience-conditioned generation is not audience testing. A bounded screen of the review’s 122 authoritative stable-key titles surfaces one explicit GPT study: Beyond Generating Code evaluates 91 quiz questions and nine homework assignments, and its one-commit supplement pins prompts, outputs, code, and grader bundles. That is a real custody upgrade. It is not an immutable provider snapshot, a versioned release, or evidence that the government, UN, or primary-school audiences named in prompts received or evaluated the artifact. There is no later commit or audience recheck. Report this as one pinned-output bundle and zero delivery/recheck chains. The other 121 title nonmatches are not exclusions, and the complete denominator remains unknown.
A mature visualization afterlife does not validate an AI feature. Exact- DOI discovery now covers 114 authoritative review keys and 87 available abstracts. It adds no hidden GenAI signal, but it surfaces AIDSVu: ten years of public delivery, broad aggregate use, named planning applications, governance, and a later 2026 data release. Those are strong platform receipts. The held surfaces expose no AI role, immutable build, version-bound measured audience outcome, or affected-audience recheck. Keep AIDSVu as a non-AI comparator. Do not borrow its reach and afterlife into a model claim, and do not turn 86 abstract cue nonmatches into exclusions.
A human study is not deployment. Primary-text recovery below the 27 no-abstract DOI rows puts five exact texts into custody: four empirical papers and one review. The empirical papers separate controlled audience measurement, public-sector co-design/demo/bounded use, and enterprise demonstrator plus production intent. None binds an accepted field release to a later affected- audience recheck, and all five explicit GenAI cue screens are empty. Keep 22 uncaptured DOI rows and eight non-DOI keys unassessed; never join the task, demo, and production-intent receipts across artifacts.
A field-use receipt is still not an AI lifecycle. The eight stable keys below A208’s DOI layer were partly an authority problem: exact publisher records recover three DOIs, correct two publication years, and show that one printed PMID names a different paper. Four full texts, two primary abstracts, one issue excerpt, and one gated primary text are now distinguished. The strongest practice receipt is a 2012 EMR cancer diary: its abstract reports increasing use in system logs, while an independent systematic review recovers an 11-clinician QUIS survey with median satisfaction of 4.38. That is meaningful embedded-use and usability evidence. It exposes no AI role, immutable accepted build, patient or decision outcome, later material event, or affected-clinician recheck. Keep the case as a non-AI operational comparator, keep the gated row unassessed, and keep the complete-review lifecycle denominator unknown.
Seven human evaluations are still not one deployed lifecycle. A targeted recovery of nine high-signal DOI rows adds eight substantive primary abstract or official-project surfaces and seven explicit participant/evaluator studies, leaving 14 DOI rows without substantive primary content. The strongest delivery near-miss is InfoViP: seven FDA safety evaluators shaped and evaluated the prototype, suggestions were addressed, and an official FDA page says an enhanced NLP and unsupervised-learning version will be installed in production. That is prospective. No held receipt confirms installation, accepted release, routine use, regulatory outcome, later change, or evaluator return. Keep the case as a non-generative production-intent comparator.
The remaining DOI content gap is one record, not a deployment finding. All 14 residual rows were investigated; eleven gained substantive primary content in the residual pass, including one full publisher chapter, while B92, B91, and B110 were then metadata-only. The chapter connects three UX experts and 25 problems to a third design version. A corridor stakeholder case and university-network case add practice context. None has an explicit GenAI visualization role, accepted field release, routine-use denominator, consequential outcome, or later affected-audience recheck.
One named interface is live; that is delivery, not observed use. A final- three recheck gives B110 an official abstract, pinned framework source, and a DiscoverWater interface reachable at KU in August 2026. At that pass, B92 and B91 remained content-unassessed. No manifest binds the live bytes to a commit, and no held receipt shows acceptance, continuous or ordinary use, consequential audience outcome, later affected-audience return, whole cost, or an AI visualization role.
Public source only partly reproduces what is live. The sole published v1.2 page state does not match the dated live page, and the linked v2.0 source is an R/Shiny application. Eight of twelve live dependency names occur in pinned application source and two of three sampled asset pairs match canonically, so partial lineage is real. It still supplies zero exact-build or measured audience-outcome receipts; a prototype demonstration and analytics hook do not change that state.
One exact publisher abstract closes one content gap, not a lifecycle. B92’s publisher record describes a multicriteria hydrogen-pipeline risk model, Monte Carlo simulation, Kendall’s tau rank comparison, and graphs for ranking sections and targeting mitigation. The abstract names no human sample or AI visualization role and supplies no release, delivery, observed-use, outcome, later recheck, or whole cost. B91 remains content-unassessed, so the current 27-row ledger is six full texts, 20 primary abstract or official-content rows, one unassessed row, and zero complete AI visualization lifecycles. Three bounded lawful passes have exhausted the same six named B91 routes, so generic retrieval is paused with four exact reopen conditions. That is one access-blocked row and zero active generic B91 targets—not a negative finding or permission to drop it.
A review corpus is not automatically a lifecycle denominator. One 2025
systematic review
maps 127 studies of visualization in broadly defined
AI-assisted decision-making: 118 empirical papers and nine reviews. It
explicitly does not focus on GenAI and does not jointly extract immutable
artifact, delivery, consequential decision or calibrated-trust outcome, and
later recheck. The complete GenAI lifecycle count is therefore unknown—not
zero of 127. Use the review for coverage and primary-source discovery, not as a
negative denominator. Its own publisher
supplement
adds a prior gate: the six domain totals sum to 127, but the “full list” names
126 citations and leaves one Education entry as a literal question mark. The
review’s own post-selection Education paragraph and bibliography now
reconstruct that row uniquely as Hernández-Calderón et al.
(2023), with matching DOI metadata. Keep
the published ? beside the repair overlay. The completed six-domain
crosswalk now binds 122 of 127 supplement positions to distinct A208 review
keys. Five printed rows have exact-looking external DOI candidates but no
review-controlled join: Jiao (2022),
Burt et al. (2017), Theis et al.
(2018), M. Lu
(2020), and Islam et al.
(2022). They remain candidates,
not authoritative keys. One admitted key also carries a printed PMID that
resolves to a different paper. The result is 122 review-owned keys, five
authority gaps, one quarantined identifier conflict, and an unknown lifecycle
count; screening has not started.
These controls are not a validated best configuration. A five-threat audit found no current Vizier configuration result to promote. Common-task studies show why: models can score strongly without reproducing human error patterns or human interpretive strategies. A benchmark therefore licenses only its named task-and-grader lane. Moving to overall quality, audience fit, comprehension, accessibility, accepted delivery, maintenance, or an architecture winner requires a new receipt at that layer. Older human task definitions may remain useful, while the exact model and harness result must be refreshed after material releases. The current qualifying configuration-run count is zero.
Natural language is becoming one control surface inside a structured environment—not a replacement for the environment.
For a governed-BI pilot, “declare context” now has a concrete minimum: name what the assistant can traverse beyond the visible report, the effective identity and permission state, the semantic objects included or excluded, the expected answer or refusal, and the owner of reproduction, correction, and regression. One captured practitioner issue and one Microsoft-authored customer account make those questions concrete, but remain C-grade self-report. They do not establish a product bug, security breach, general efficacy, prevalence, or demand. See the audience decision report.
Capability improved quickly, but the target also became harder
From 2023 through August 2026, general models gained image understanding, code execution, tool use, browser operation, persistent workspaces, and reusable skills. Research advanced from single-chart generation and chart-to-code tasks to editable multi-view dashboards, replayable interaction, realistic chart reading, visual integrity, and expert-adjudicated critique.
Within individual benchmarks, later systems often make large gains in visual fidelity and task breadth. Execution and semantic reliability improve less uniformly. Across benchmarks, there is no honest single progress curve because the inputs, models, renderers, judges, and definitions of success change. Stronger baselines also expose harder questions: not merely whether the chart renders, but whether the data binding is right, the interaction works, a repair introduces a regression, or a reader learns the intended point.
Two August benchmarks sharpen that boundary. Dashboard2Code now tests stateful reconstruction across 180 Plotly Dash applications and 450 interaction tasks; hidden state and transformations can be wrong while the interface responds. Chartography’s best tested configuration reaches 45% mean pass@1 on 100 deliberately difficult, practitioner-authored professional chart-reading tasks. The latter is not an ordinary-chart failure rate. Together they show that interactive state and professional visual conventions are distinct capability frontiers, not details covered by a generic screenshot score.
DV-World adds the next pre-delivery rung without proving production maintenance. Across 50 native Excel chart repairs, the best reported agent route succeeded 48% of the time; across 80 changed-data or evolving-requirement tasks, the top overall score was 51.44%. Those tasks make binding failures and destructive regressions testable. They still end at one benchmark-scored candidate: no deployed artifact, elapsed later event, maintainer, reader outcome, direct-work baseline, or whole cost was observed.
Three same-artifact production afterlives now exist, with different evidence strengths. A four-month KubeStellar Console experience report, commit-pinned quality record, and linked pull requests follow one AI-assisted dashboard codebase through a live runtime regression, repair, a separate build blocker, restored deployment, and a manually added prevention check. That is a real rung beyond DV-World’s prepared changes. It is one solo-maintainer project-authored case, not an independent productivity comparison or evidence of safe autonomy.
OpenClaw Agent Dashboard
adds a second crossing and the first here with a non-owner operator report. The
project says it was built with Claude Code and released v3.0.0 on March 5,
2026. Thirteen days later, an operator reported a cost view showing
$0 for a
custom provider despite present token data; the owner acknowledged that the
configuration had not been tested. The captured issue remains open. This is
distribution-to-failure evidence. A different non-owner has since opened
PR 31, which
proposes a provider-inference fallback matching the reported mechanism. The
proposal has no maintainer review, merge, corrected release, or reporter
recheck, so it is not recovery. Project-level AI attribution does not identify
who wrote the defective line.
Prism adds the first held multi-attempt delivery-and-repair chain. Its maintainer says the family dashboard was built entirely with Claude Code under human product direction. In issue 81, a non-owner operator reported an initial Home Assistant add-on install failure, a failed repair build, and then install success followed by startup failure; the operator used ordinary Docker instead. The later v1.8.11 release and release-pinned entrypoint repair the recorded socket and schema-replay failures, and the maintainer reports install, start, and restart checks on real Home Assistant OS. The affected operator never rechecked that release. This is repaired-release and maintainer-verification evidence, not independent recovery, and it does not attribute any defect to Claude Code.
The pull-request histories add a separate, bounded form of handoff evidence. Prism’s non-owner PR 23 carried a Claude-assisted weather visualization through maintainer-found delivery defects, contributor repairs, owner integration, and a credited release; the contributor returned with accepted PR 41. OpenClaw’s non-owner PR 15 put Claude-assisted multi-provider pricing into the product after maintainer integration. These are receipts for accepted second-person change, and Prism adds repeat contribution. They do not name a second maintainer, transfer release or incident authority, or independently accept either later recovery.
The consolidated production result is short: three afterlives, two maintainer-verified restorations, zero affected-user or independent-operator recoveries, zero transferred maintenance authority, and zero complete twelve- state cases. A closed issue is not a recovery receipt, a maintainer check is not an affected-user check, and an accepted contribution is not operational authority. The next evidence worth funding is a return-user exercise of the corrected route or a named second person actually releasing or responding to an incident.
A second single-case PM4Py-UCM experience report measures about 65 active hours across 18 agent sessions and ten weeks; fixes outnumbered features 2.3 to 1, and 78% of visualization and layout work was fixes. It makes refinement cost visible without following a later field incident or reader. Neither case supplies the complete human-and-model cost, maintenance-authority handoff, accessible-reader use, consequential decisions, or calibrated trust needed to close the lifecycle question.
A project history is not a same-artifact afterlife. A closed audit of four promising public candidates admitted zero new episodes. The strongest near-miss is MAIDR: 11 blind participants tested four chart types using their own screen-reader and refreshable-braille-display paths, but the studied implementation predates AI. The current project declares a complete TypeScript rewrite and an AI description layer; no captured source pins a representative participant recheck to that release. Separately, a public-health copilot paper describes a 16-person trust/usability method but says the experiment was removed and reports no human result. Methods, results, release, and later use remain separate gates.
The public legacy code now narrows the later AI study further without closing
that gate. A dedicated box-plot study surface is present before
v2.10.0, making that
tag the first containing version floor found—not proof of the build or
model used by the eight blind and low-vision participants named in the
abstract. Nine selected later model, verification, and deprecation events add
real maintenance history but no representative-user recheck. The separate
TypeScript rewrite needs its own validation.
The executive decision rule is: preserve the exact artifact from acceptance through deploy, incident, repair, corrected release, independent recheck, restored service, and the new regression gate—then count the human infrastructure and follow the result to a second maintainer and intended readers. Track accepted change, repeated contribution, maintenance authority, and independent recovery as separate states. Treat unknown, zero, and not applicable as different states.
The benchmark crosswalk shows exactly what each major evaluation measures and cannot establish. The capability scorecard turns that evidence into current, historical, and forecast scores by use context.
The forecast derived from this record is deliberately asymmetric. It assigns higher near-term probabilities to accepted static work, interactive-dashboard reasoning, narrow visual critics, and environment-specific adapters because they have public tests and observable failure signals. It assigns lower probabilities to production productivity, representative-reader benefit, and autonomous publication because the evidence needed to verify those outcomes is mostly absent. Better models can help request and reason over local definitions, provenance, authority, and reader response; they cannot manufacture those facts.
The first forecast check resolves nothing—the earliest date is August 2027—but it does correct the record. ViviDoc was published before the forecast freeze and captured afterward, so it is disclosed as a pre-freeze source omission rather than a new capability event. Its typed State–Render–Transition–Constraint plan improved same-pipeline interactive-document measures across three backbones. It did not measure representative readers, accessible delivery, whole production cost, later maintenance, or autonomous publication, so all seven forecasts remain open and unchanged.
What consistently helps
- Real context. Field definitions, measures, joins, units, permissions, source vintage, audience, and delivery constraints prevent errors a generic persona cannot.
- Inspectable representations. A semantic chart object, analytic specification, restricted language, or visible intermediate table makes validation and local correction easier. Direct code remains appropriate where expressiveness matters and the risks are low.
- Mixed-initiative control. Natural language works better beside direct manipulation, visible transformed data, history, branching, and reversion than as an endless regenerate button.
- Layered verification. Query execution, recalculation, schema checks, rendered inspection, browser replay, accessibility checks, and human review establish different facts. None substitutes for the others.
- Bounded specialization. A separate agent, skill, critic, or vision model helps when it owns different information, a tool, or an evidence channel. A new job title alone does not create specialization.
These are shared principles, not a single architecture. Journalism, governed BI, operational monitoring, exploration, science, education, and reusable applications have different purposes, authorities, update cycles, and failure costs. A common record of intent, data, transformations, artifact, revisions, and evidence should route to different authoring and acceptance procedures.
The specialist-cost question is now measurable, but not measured. Nine primary studies report fragments: roughly 4–16× inference for repeated chart extraction; about 31–33 seconds of human verification per chart in ExChart; about 6–10 seconds for one GPT-4o call versus 90 seconds serial or 30 seconds parallel and about $0.40 per query for ChartAgent; substantial crowd and expert work in VisJudge-Bench; and 450 manually annotated dashboard interactions plus human validation in DashboardMimic. A sixth chart-specific routing paper separates training resources from per-chart inference, but its displayed split totals, latency percentage, memory-delta field, and selected configuration do not reconcile; its exact numbers are quarantined. METAL adds same-task quality results, but its five-candidate baseline makes five generation calls while its five-recurrence multi-agent route can make one initial generation plus fifteen critique or revision calls. Neither its paper nor released logs preserve the per-arm usage needed to reconcile that difference. A new ChartAgent tool-integrated-reasoning preprint reports same-task accuracy beside 3.0–9.8 mean tool invocations across scheduler settings. The count omits model and local-tool compute, reflection, tokens, latency, retries, evaluator and human work, and task-level accepted outcomes. None of those eight visualization fragments joins one current-general and specialist route through acceptance, delivery, readers, and maintenance, so they cannot be added into a route total.
The ninth fragment is adjacent rather than visualization-specific. ProMCP shows measured token and latency bottlenecks moving between planning/schema injection and final synthesis as MCP client topology changes. Its pinned release has useful event fields but no task traces or historical run manifest, and its benchmark entry point is not runnable as released. The result adds cold start, discovery, phase, topology, cache, streaming, concurrency, retry, and observability state to the receipt; it does not add a visualization baseline or cost winner.
The decision rule is an eleven-lane trial: start with route preparation and ownership, then keep task/acceptance, success, inference, latency, evaluation, human correction, unresolved failures, maintenance, delivery, and reader outcomes separate under one equal total budget. Freeze a multidimensional cap, then retain observed calls and retries, input/output tokens or image units, evaluator and tool work, accelerator use, latency and concurrency, charges, and human minutes for every task. Matching rounds, candidates, rates, nominal caps, or model labels is not an equal-budget receipt. Distinguish fixed preparation, periodic ownership, marginal attempts, failure-contingent work, and downstream delivery or reader cost; declare the volume and time horizon before amortizing. Within machine work, retain initialization and discovery separately from planning, tool calls and results, context updates, and answer synthesis; bind each event to the effective host/client/server/model/tool topology and runtime state. Hidden internal phases stay missing rather than inferred. Recompute source denominators and percentages before admitting telemetry, then report cost per eligible, attempted, accepted, and reader-successful task. This turns the gap into an executable procurement and research question without claiming that a specialist or general route is cheaper.
Cost per accepted artifact now has an explicit two-receipt rule. Use all reconciled route cost over the declared population and observation window as the numerator, including rejected, abandoned, quarantined, and no-output work. Use only artifacts passing the frozen human or production acceptance contract as the denominator. Tool calls, per-attempt API averages, aggregate accuracy or F1, finish signals, and one-arm telemetry are not substitutes. Zero of the twelve held fragments supplies both receipts.
One result now clears a narrower bar. The peer-reviewed Selective TTS visual- insights study matches a declared LLM-call budget and, in a separate experiment, output tokens inside one pipeline. Its released code still omits repair and errored-worker cost, records completion tokens only, and counts proxy-judge-scored candidates rather than contract-accepted artifacts. Treat this as a matched declared partial inference budget—not observed total use or cost per accepted work.
A second result makes retry burden visible without pricing it. On 888 VisPlotBench tasks, VisCoder2-32B begins with 649 execution passes and GPT-4.1 with 563; after up to three conditional self-debug rounds, both finish at 732, or 82.4%. The paper tables reconstruct to 584 versus 714 revisions, with diminishing rescue in later rounds. The released debug paths do not retain comparable task-level tokens, compute, latency, or charges, and execution pass is not human or production acceptance. Read this as same endpoint, different retry topology, no cost winner—and persist resource telemetry before using a retry count in procurement.
What creators and readers actually experience
The practitioner and reader experience explains why enthusiasm and frustration coexist. First drafts can appear before an idea has cooled, and tedious implementation can compress dramatically. Correction and verification can erase that gain. In one study of 18 experienced analysts completing 108 deliberately error-prone tasks, seven episodes were unfinished and participants prematurely accepted the result 31 times. A novice study found low verification intent and repeated failed repairs alongside positive satisfaction.
Expertise changes both the benefit and the burden. Experienced practitioners can filter weak suggestions and use AI selectively, but they also notice when describing a small edit takes longer than making it. In interviews with 17 biomedical-visualization practitioners, AI was used chiefly for auxiliary work; all opposed substantial generated visuals in final scientific communication.
Expertise is better represented as a capability profile than a rank. Established visualization-literacy research distinguishes consuming, constructing, critiquing, and connecting a chart to its context; professional work adds data, domain, implementation, situated-judgment, and delivery resources. Current learner studies bank access, reported speed, confidence, and engagement more readily than correctness, verification, or delayed independent skill. One 117-person randomized visual-comprehension study found an immediate post-removal advantage for a proactive agent that asked scaffolded questions; it did not test delayed construction or far transfer. Adjacent randomized education studies show why the distinction matters: assisted exercise performance can improve while conceptual learning does not, and unrestricted answer access can harm later unassisted work. Small studies suggest intermediate and expert practitioners can turn structured critique and constrained implementation into better work more reliably, but no large study establishes one universal expertise curve.
The decisive learning study is now concrete, but still unrun. It would compare conventional instruction, answer-oriented AI, and metacognitive AI using the same frozen model and harness; remove assistance immediately, at six weeks, and at six months on unfamiliar tasks; then test the resulting artifacts with blinded readers on their own devices and assistive paths. Ten outcome lanes keep access, assisted performance, learning, cost, accessibility, and reader use separate. No owner, ethics approval, powered sample, participant, budget, or result exists, so this design narrows the research gap without answering it.
Readers receive the claim, not the authoring transcript. In a controlled 48-person study, selected AI-generated misleading charts reduced answer accuracy from 88.3% to 71.9%. Adjacent accessibility studies show why delivery cannot be inferred from a responsive screenshot or text alternative: outcomes for low-vision and blind readers changed with the actual device, interaction, modality, chart type, time, and workload. A new 12-person blind and low-vision study adds direct AI-assisted learning evidence: eleven preferred tactile charts plus text and chat and described a stronger spatial model, but measured chart-understanding accuracy did not improve. No current study joins AI generation to a representative mobile or assistive-technology reader evaluation.
Following the artifact past its creator exposes another class of failure. In one public second-maintainer episode, an AI-built dashboard failed during the creator’s absence because refresh and scheduling lived on that person’s laptop; the inheritor replaced it with a proper pipeline. This is testimony, not a prevalence estimate, but it makes execution host, lineage, definitions, dependencies, documentation, ownership, and rollback part of acceptance—not optional handoff cleanup.
Skills can help; popularity does not say whether they do
The agent-skill deep dive finds that the public ecosystem contains hundreds of listings labeled as data visualization, with extensive copying, vendoring, and source drift. Registry installs measure acquisition. Repository stars usually belong to a much larger repository. Neither measures routine use or output quality.
The inspected packages perform six different jobs: primers, storytelling guides, renderer gateways, library or domain adapters, quality gates, and structured workflows. Their most defensible value is information a capable model cannot reliably infer: a current API, local environment, semantic representation, known failure, executable validator, or delivered-surface repair loop.
There is direct evidence that this can work. SciVisAgentSkills improved quality in all ten paired suite-by-agent comparisons across 108 scientific- visualization tasks, although one completion measure fell. Broader benchmarks show the boundary: compact, relevant, compatible skills can help; comprehensive, self-generated, stale, or incorrectly retrieved guidance can perform worse than no skill. Popular generic, dashboard, accessibility, mobile, and explanatory packages still lack independent paired evaluation.
Specialized vision is a sensor, not a final judge
The specialized-vision deep dive finds that “vision for charts” includes at least six jobs: question answering, structured extraction, OCR and layout, element grounding, integrity or perceptual critique, and verification or repair. No model covers all six reliably.
Specialists show real component gains. VisJudge-7B matched its expert- adjudicated quality rubric better than the tested general models. ChartAgent’s chart-specific tools materially improved the same general reasoner on numeric questions. Chart-aware grounding and recent OCR/layout models improve bounded perception. At the same time, newer out-of-distribution extraction and realistic chart-QA tests show strong general models overtaking older chart specialists. Dashboard2Code adds a stateful fixed-desktop test but leaves responsive/mobile, animation, keyboard, and assistive-technology behavior open. Chartography shows that domain conventions and hard professional visual forms still defeat strong general systems even with more reasoning.
The practical choice is role-based: use source data and deterministic checks whenever available; add a specialist for a measured parsing, localization, or perceptual failure; retain a current general model for broad reasoning; and do not confuse any model’s score with reader comprehension or publication acceptance.
What performs poorly, adds risk, or remains unsupported
- Proxy as proof: treating a nonblank render, attractive screenshot, code-similarity score, or one model judge as proof of correctness;
- Unbounded repair: unlimited self-critique or repair without a bounded fault, budget, regression check, and stopping rule;
- Role-play instead of specialization: multiplying agents or personas without distinct information and tool boundaries;
- Synthetic readers: using personas as representative readers or domain experts;
- Popularity as evidence: assuming a skill helps because it is long, popular, installed, or official;
- Specialists outside their lane: applying a chart specialist outside the distribution and role it was tested on;
- Premature storytelling: forcing a narrative arc before the analysis supports one;
- Desktop shrinkage: shrinking a desktop composition and calling the result mobile; and
- First-render accounting: claiming production value from first-render speed without counting prompting, waiting, verification, failed repair, deployment, abandonment, and later maintenance.
The next evidence should follow work to acceptance
The capability forecast and research-package map show where evidence is still missing. The largest gaps sit between measured capability and lived outcome:
A ten-case bridge audit narrows that statement without reversing it. Three controlled studies connect technical evidence to analyst use, analyst correction, or reader harm. Zero audited case binds an exact system version and frozen acceptance rule to real delivery, and zero follows one artifact through reader or decision use, a later maintenance event, and whole cost. The bridge is short, not empty—and it is not a complete lifecycle.
- Acceptance: the share of ordinary work accepted without repair;
- Semantics: fidelity against owned measures and changing sources;
- Correction: cost, regression, stopping, and abandonment;
- Routing: equal-budget comparisons among deterministic checks, general critics, specialist models, and skills;
- Who gains and how: capability-profiled gains across access, productivity, quality, learning, verification, reader outcome, and sustained work;
- Learning: delayed unassisted tests that distinguish learning from dependence or skill atrophy;
- Operations: authenticated delivery, multi-context handoff, maintenance, and total cost;
- Governance: organizational ownership, privacy, provenance, disclosure, and recovery after error; and
- Reader outcomes: comprehension, decisions, calibrated trust, accessibility, and mobile use by the intended readers.
One current readiness audit makes the experimental gate concrete. It reconciled 14 historical case definitions, 18 run directories, and 13 judge directories, and now has all five required context slots with executable source packets. The exploratory packet is a commit-pinned, CC0 Palmer Penguins exploratory fixture: across 342 complete rows, bill length and depth correlate -0.235 when pooled but positively within every species. That reversal can test source, grouping, branching, and render receipts. Because the dataset is familiar and public, it is not held-out efficacy evidence or a human outcome. The third packet is a first-party synthetic governed-BI contract: an owned measure, vintage, four roles, six verified queries, four refusal cases, and an exact fanout incident. A bad tag join returns $16,830 against the authoritative $7,730. That makes semantics, permission, refusal, and recovery testable—not production security or reliability. The fourth packet uses official 2024 ACS B19013 estimates and 90% margins of error. Iowa, Kansas, Montana, and Wyoming differ by only $13 to $192; no internal pair clears the declared 90% comparison threshold. That makes exact uncertainty propagation, refusal of a false winner, and expert review routing testable—not a scientific result across domains or an expert or reader outcome. The fifth packet is a first-party synthetic interactive maintenance explorer with two data versions, one canonical state, ten expected states, seven isolated trajectories, keyboard and narrow-delivery contracts, and a stale-selection failure. When a filter makes T007 ineligible and leaves only T012 visible, detail, accessible summary, URL, and export must clear T007 together. That makes synchronization, replay, recovery, and prepared change testable—not a browser run, accessibility result, human outcome, or production maintenance result. Historical auditability and five-of-five source-packet readiness are not a current baseline: compare no architecture until the exact current model/harness manifest is frozen and every arm produces equivalent receipts.
A first current-harness preflight made that distinction executable. Its seven-blocker predecessor locked the five source packets, local toolchain and available-skill observations, task order, three arm roles, and an 18-field common receipt shape. A credential-free neutral runner now clears one of those blockers: it recomputes five source checks, stages 15 isolated fixture-arm attempts, separates model input from evaluator-only truth, hashes artifact, render, trace, and check files outside the arm, and rejects missing or mismatched receipts. This matters because the legacy editorial case embeds its answer key; handing over the whole packet would leak the evaluator into the task.
The successor preflight still reports six pre-run blockers, zero model calls, and no spend authority: the exact served-model snapshot and server harness are not exportable, no same-model bare interface is available, no one Vizier mechanism or acceptance authority is selected, and no capped model budget is authorized. The passing runner self-test is evaluation plumbing—not a model, artifact, browser, human, or architecture result. A model label and local CLI version remain useful inventory, not a reproducible comparison.
The proposed Vizier delta is now narrow enough to decide without calling it selected. The two public skills each combine framing, retrieval, guidance, tools, inspection, and critique; either whole package would be another bundle comparison. The prepared candidate adds one artifact-evidence reconciliation checkpoint after the first executable artifact and before repair or finalization. It asks the arm to bind visible claims and states to available source, transform, query, calculation, branch, or canonical-state evidence and freeze mismatches before repair. Its corrected schema refuses an overall pass when any surface is mismatched or unresolved, required states are missing, or repair remains necessary. It adds no model, agent, tool, answer key, or call. The Vizier owner has not ratified it, so the layer remains null and the preflight still has six blockers. A testable protocol is not an intervention choice or an effect result.
The acceptance gate is also narrow enough to decide without pretending a model
or checklist is a person. One candidate separates deterministic checks from six
nonmechanical lanes: context review, domain expertise, accessibility
conformance, evaluation with relevant disabled users, intended-user acceptance,
and later maintenance handoff. Across five fixtures, a mechanical-only pilot
would retain 30 explicit not-run records. A human pass requires a named
eligible human authority and evidence; an automated tool cannot occupy that
role. W3C’s evaluation overview says no
tool alone determines accessibility, while its
user-evaluation guidance
keeps user experience distinct from conformance and warns against generalizing
beyond the participant scope. No acceptance scope is selected and no person was
contacted, so the research-owner gate remains open and the blocker count stays
at six.
Those six blockers now have one attributable response contract rather than one blanket “go” decision. The register exposes 17 admissible paths across model- provider/runner, harness-provider/runner, research-owner, and Vizier-owner authority. It rejects wrong-role, incomplete, secret-bearing, partial, and six-answer-without-release states. Provider facts cannot authorize spend; credential availability cannot amend a protocol; a passing candidate verifier cannot select a layer or acceptance scope. Even six advancing answers require a separate research-owner release of the resolved decision-set digest, named run plan, call cap, and stop condition. All six selections and release remain null, readiness and spend remain false, and calls and contacts remain zero. The research review translates that same custody state for practitioners, BI leaders, researchers, newsrooms, product teams, educators and accessibility specialists, and broad AI readers.
Future receipts also have a scorer that cannot hide missing evidence inside one
grade. It emits computable and checkable mechanical tiers plus six named
context, domain, accessibility, user, and maintenance lanes. Empty or not-run
mechanical evidence remains incomplete; human authority is bound to the exact
lane; overall_score and universal_pass are invalid fields. Its verifier
created 15 synthetic fixture-arm ledgers and preserved 90 human not-run
states with zero calls. Those are plumbing tests, not successful AI artifacts.
The research review
shows how practitioners, leaders, researchers, newsrooms, product teams,
educators and accessibility specialists, and broad AI readers should interpret
the same lane-specific record.
The next cheap hypothesis test is also preregistered but unrun. R4 now has five opaque, hashed visualization inputs: a 19-file model allowlist keeps five visuals, five briefs, and nine source-evidence files separate from five hidden defect records. Two hashed prompts and two fail-closed schemas keep integrity and readability separate; they accept findings, empty completions, and abstentions while twelve adverse records reject aggregate, mixed, identity, evidence, and state failures. It would compare three exact models with one independent human data editor and a predeclared tie rule. Model slots, editor calibration, budget, release, and results remain absent. Five cases are a protocol pilot, not confirmation; no ranking or divergence exists.
The five residual gate categories now have a seven-receipt decision surface: three exact served-model receipts, editor acceptance, same-editor calibration, a capped 30-primary-call budget, and final release over the exact prerequisite digest. Twenty advancing and non-advancing paths validate and sixteen adverse states fail closed. This makes the handoff auditable; it supplies no real provider, editor, spend, release, call, ranking, or result.
The editor-calibration handoff is also concrete but blank: ten baselines, ten synthetic response-quality vignettes with two swapped repeats, three valid response states, and fourteen rejected adverse records. It prevents model output from setting the baseline, but it does not appoint an editor or produce a resolution, tie threshold, score, ranking, or result.
The next experiments should begin with the current model and its normal harness, use representative tasks from declared environments, separate first draft from accepted artifact, add one mechanism at a time, exercise the delivered surface, and retain every correction and regression. Additional instructions, agents, or models should survive only when they improve that full path for a stated job at a proportionate total cost.
Update log
-
2026-08-16 — Turned the final content gap into a stop-and-reopen decision. B91 stays access-blocked and content-unassessed in the unchanged
6 / 20 / 1 / 0ledger. After three bounded passes, generic search is paused until exact new lawful custody appears; effort can move to a stronger lifecycle-bearing target without manufacturing a negative row. -
2026-08-16 — Mapped InfoViP’s research contacts without inventing an operating owner. The closed Elsa/API/UI opportunity names three research mentors but no selection or execution; their official profiles support research and project roles, while operation, maintenance, QA execution, release approval, authorization, and the application join remain unassigned.
-
2026-08-16 — Repaired InfoViP’s 2024–2026 episode dates and relationships. An apparently 2026 future-tense FDA page is explicitly content-current to 2024, so it does not contradict the 2025 component report. A 2026 FDA biography adds a current project lead without operations authority; the Elsa 4.0/HALO launch adds no InfoViP-specific integration. The correction changes the decision trace, not the 15 named / 0 complete count.
-
2026-08-16 — Final InfoViP evidence upgrades the component, not the whole lifecycle. A final CIOMS report says the deduplication component was approved, installed in AWS, integrated with AERS, and processing more than 30 million historical plus about 8,000 daily submissions. The same report keeps the QA plan, completed audits, routine roles, and downstream effects open; the current inventory still records implementation N/A and no ATO. A 2026 Elsa/API/UI opportunity is a prospective recheck trigger. The audit remains 15 named cases and zero complete episodes.
-
2026-08-16 — InfoViP adds counted field use and a bounded award envelope, not a complete lifecycle. Twenty unique reviewers made 58 submissions in a six-month internal real-work evaluation; a later primary abstract reports a 29-million-history plus daily operating pipeline; three related FDA awards total $2.40 million obligated and $2.86 million estimated. Current implementation/ATO, routine post-release use, outcome, return, and whole cost remain unresolved; 15 named cases still yield zero complete episodes.
-
2026-08-16 — InfoViP adds an operational pipeline, not a complete deployment. Joined seven-evaluator co-design to an operational pilot, 28-million-report plus incoming-report processing, a human-in-the-loop performance report, and current governance counterevidence. The audit now contains 15 named cases and zero complete episodes.
-
2026-08-16 — One publisher abstract reduced the content gap to one. B92 adds an abstract-level non-AI uncertainty-and-ranking method, not a human-use or delivery lifecycle. The current ledger is 6 full / 20 abstract or official / 1 unassessed / 0 complete AI lifecycles; B91 remains unassessed.
-
2026-08-16 — Public source matched partially, not as a deployed build. The live DiscoverWater page differs from the sole published page state; two of three sampled data assets match canonically. No exact-build, audience-use, outcome, accessibility, whole-cost, or AI receipt was added.
-
2026-08-16 — One public-delivery afterlife reduced the gap to two. B110 gained an official abstract, pinned primary source, and a same-named live KU interface. At that pass B92 and B91 remained unassessed; no exact live-build join, use denominator, audience outcome, recheck, whole cost, or AI role was added.
-
2026-08-16 — Residual DOI content gap reduced to three. Investigated all 14 residual rows, recovered substantive content for eleven including one full chapter, and retained three metadata-only rows. Added an expert evaluation-to-redesign mechanism and two practice-context near misses without admitting a new AI lifecycle.
-
2026-08-16 — Seven evaluations kept below deployment. Eight new substantive primary abstract or official-project surfaces add seven human studies and one InfoViP production-intent near-miss, but no confirmed installation, accepted release, outcome, or later recheck. Fourteen DOI rows remain content-unassessed.
-
2026-08-16 — A field-use receipt kept below AI lifecycle. Repaired the eight-key non-DOI stratum and added one non-AI embedded-use/usability comparator without promoting it to an accepted AI release, outcome, or later audience recheck.
-
2026-08-16 — Human evidence kept below deployment. Five recovered full texts add four empirical pre-deployment cases across controlled task, public-sector co-design/demo, and enterprise production-intent lanes. None reaches accepted field release plus later affected-audience recheck; 22 DOI rows and eight non-DOI keys remain unassessed.
-
2026-08-16 — Mature delivery kept separate from AI efficacy. AIDSVu adds ten years of public delivery, aggregate use, governance, and a later data release to the review screen, but no AI role or version-bound affected- audience recheck. It remains a non-AI comparator; 86 abstract cue nonmatches remain unexcluded.
-
2026-08-16 — Audience-conditioned prompts kept below audience testing. One explicit GPT title among 122 authoritative A208 keys leads to a pinned one-commit output bundle, but no versioned release, delivered intended audience, outcome, or later recheck. The 121 title nonmatches remain unscreened, so the complete denominator stays unknown.
-
2026-08-16 — Development selection kept below audience impact. Lexara adds a two-week field-deployed evaluation near-miss with six developers, 38 experiments, 57 newly authored cases, ten models, and six prompts. Across 14 named reader, production, and evaluation cases, zero joins immutable build, accepted delivery, consequential decision or calibrated trust, whole cost, and later same-lineage recheck.
-
2026-08-16 — Full review frame stops at 122 of 127 authoritative keys. Five exact-looking DOI identities lack review-controlled joins, one printed PMID points to another paper, and lifecycle screening remains unstarted. The executive boundary keeps identity confidence, identifier validity, and admission authority separate.
-
2026-08-16 — Education key audit finds one unresolved authority join. Eleven printed labels plus B59 yield twelve A208 DOI keys.
M. Lu (2020)has a plausible Crossref candidate but no A208 bibliography record, so the executive boundary reports 12 of 13 authoritative Education keys and keeps lifecycle screening unstarted. -
2026-08-16 — Missing review row repaired without upgrading the denominator. A208’s article uniquely reconstructs the Education
?as Hernández-Calderón et al. (2023), DOI10.1093/iwc/iwac043. The executive boundary preserves the published source value, repair provenance, absent publisher correction, unfinished stable-key crosswalk, and null lifecycle count as separate states. -
2026-08-16 — Reported review count separated from row identity. A208’s appendix preserves 127 domain rows but names 126 citations and one unresolved Education entry. The executive gate now requires inventory repair before a complete primary-paper screen; neither the count nor the unknown lifecycle result is silently rewritten.
-
2026-08-16 — Review corpus kept below lifecycle denominator. A five-database review establishes breadth across 127 decision-visualization studies, but its non-GenAI focus and unjoined lifecycle fields keep the qualifying count unknown rather than zero of 127.
-
2026-08-16 — Retry topology kept below accepted cost. Across 888 tasks, VisCoder2-32B and GPT-4.1 both finish at 732 execution passes after 584 versus 714 revisions. Missing comparable resource telemetry and contract acceptance keep the result out of cost and route-selection claims; the held total is now 0/12 accepted-cost comparisons.
-
2026-08-16 — Production recovery kept actor-separated. Three held dashboard afterlives yield two maintainer-verified restorations, zero affected-actor recoveries, zero authority transfers, and zero complete twelve-state cases; the next valid gate is a return user or exercised second- person authority.
-
2026-08-15 — Human evaluation split by AI role. Added the bounded ten-row reader-outcome result: one direct controlled AI-created-chart effect, five adjacent human-outcome rows, three creator/co-design rows, one model- reader-only row, and zero complete delivered-reader recheck chains.
-
2026-08-15 — Governance claims now require one deployment row. Three provider control contracts, two named-feature organizational-use reports, and two adjacent governance processes are retained as distinct evidence; zero of seven clears the complete nine-receipt join.
-
2026-08-15 — Public acquisition is now below adoption. Twelve named skill listings collapse to ten parent repositories and eleven documented lineages. All twelve carry install signals; zero exposes public invocation, retention, organizational acceptance, or outcome receipts. Popularity can prioritize inspection, not authorize procurement or efficacy claims.
- 2026-08-15 — System anatomy now assigns authority. Across six primary authoring or analysis systems, generation-execution-critique coverage is 6/6, meaningful human control is 4/6, and accepted-delivery, intended-reader, and maintenance chains are 0/6. The result supports an eleven-stage custody map, not one universal agent topology.
- 2026-08-15 — Three controlled bridges stop before delivery. Ten primary cases yield three partial technical-to-human joins, zero exact-version accepted-delivery joins, and zero complete reader/later-maintenance/whole-cost episodes. The result narrows the capability-to-practice gap without claiming ordinary field efficacy.
- 2026-08-15 — Fixed-budget result kept below total cost. Selective TTS adds a matched declared partial call/output-token budget, while a pinned code audit keeps repair, failed-work, input/image/local compute, people, and acceptance visibly missing. The accepted-cost result is now 0/11 with no route winner.
- 2026-08-15 — Accepted cost needs a route numerator and pass denominator. Added ChartAgent’s same-task accuracy/tool-call sweep, found zero of nine held fragments with equivalent route cost plus frozen accepted outcomes, and advanced the cost protocol to ledger v5 without selecting a winner at that evidence cut.
- 2026-08-15 — A release floor is not participant exposure. MAIDR’s public
code pins a legacy AI-study surface and first containing
v2.10.0tag, while leaving the exact tested build and model unknown. Later maintenance and the TypeScript rewrite add zero representative-user rechecks. - 2026-08-15 — Exact version required for lifecycle evidence. Four public candidates yielded zero new episodes. A representative pre-AI accessibility study and a current AI-enabled rewrite remain separate until the exact release receives a representative recheck; a described-but-removed human study is not a trust result.
- 2026-08-15 — No benchmark-to-configuration shortcut. Added two direct human/model chart studies and a five-threat validity audit. Task success, human-like errors, analytical intent, reader experience, and configuration value remain separate; no current Vizier configuration winner exists.
- 2026-08-15 — Machine totals need phase and topology. Added adjacent ProMCP profiling plus a release audit; ledger v4 now preserves cold start, discovery, execution phase, effective topology and runtime state without claiming a visualization baseline, reproduced run, or cost winner.
- 2026-08-15 — Contribution is not operational handoff. Added accepted non-owner AI-assisted changes in Prism and OpenClaw, Prism’s repeat contributor, and OpenClaw’s open repair proposal. These improve evidence that another person can change the artifacts; they do not establish maintenance authority, a merged repair, or independent recovery.
- 2026-08-15 — A repaired release is not independent recovery. Added Prism as the third same-artifact dashboard afterlife and the first held multi-attempt delivery chain. The case reaches v1.8.11 and a maintainer runtime check, but no independent final recheck, whole cost, handoff, accessible-reader use, decision, or trust outcome.
- 2026-08-15 — Equal rounds are not equal resources. Added METAL as a seventh cost fragment: its same-task quality comparison is useful, but its routes have different call topologies and lack reconciled per-arm usage. The ledger now requires both a frozen multidimensional cap and observed receipts; no route ran and no cost winner exists.
- 2026-08-15 — A second same-artifact afterlife. Added OpenClaw Agent Dashboard’s release-to-failure chain and independent operator report; kept accepted repair, corrected release, recheck, whole cost, handoff, accessibility, decisions, trust, and code-level AI causality explicitly open.
- 2026-08-15 — Governed-BI context envelope added. Turned two bounded Fabric Community captures into a five-part pilot receipt for traversal, permissions, semantic inclusion, acceptance state, and correction/incident ownership. The sources remain C-grade self-report, not a platform verdict or demand signal.
- 2026-08-15 — Specialist cost now begins before inference. Added one peer-reviewed preparation/inference fragment, quarantined its exact figures after four internal reconciliation failures, and expanded the comparison to eleven lanes, five cost classes, declared amortization, and a source- integrity preflight. No route ran and no cost winner exists.
- 2026-08-15 — Specialist total-cost ledger designed, not run. Preserved five incompatible evidence fragments without summing them, then defined one same-task, equal-budget, ten-lane comparison through acceptance, delivery, readers, and maintenance. No total, route winner, or budget authorization exists.
- 2026-08-15 — Durable-learning test specified, not run. Added the three-arm, five-wave longitudinal design, separate blinded reader stage, and explicit human, accessibility, power, model, budget, analysis, and release boundaries without claiming a learning or atrophy result.
- 2026-08-15 — Forecast record corrected without hindsight rewriting. Added the ViviDoc typed-interaction result, classified it as pre-freeze evidence captured after freeze, and preserved all seven unresolved forecast probabilities and higher-level reader/lifecycle boundaries.
- 2026-08-15 — R4 editor task prepared; no editor result. Added ten baselines, ten calibration vignettes with two swapped repeats, three response states, and fourteen adverse rejections. Editor, appointment, calibration, tie, call, contact, ranking, and result remain absent.
- 2026-08-15 — R4 approval path typed; all receipts pending. Added seven attributable receipts across five gate categories, a 30-primary-call budget boundary, exact-set release, 20 admissible paths, and 16 rejected adverse states. No provider, editor, budget, release, call, contact, or result exists.
- 2026-08-15 — R4 response contract frozen; no model answer. Added two lens-specific prompts, two fail-closed schemas, three honest response states, and twelve rejected adverse records. Exact models, editor calibration, budget, release, calls, contacts, rankings, and results remain absent.
- 2026-08-15 — R4 inputs frozen; authority still absent. Added five hashed SVG stimuli, 19 allowlisted model-input files, five evaluator-only truth records, nine origin bindings, and adverse leakage/hash checks. No provider, editor, spend, release, call, ranking, or result exists.
- 2026-08-15 — R4 protocol frozen; Q8 stays open. Added a blocked five-context, three-model, two-lens preregistration with editor, tie, budget, release, and pilot boundaries. No call, contact, ranking, or result exists.
- 2026-08-15 — Scoring stays typed and non-aggregate. Added two mechanical tiers, six nonmechanical lanes, authority binding, 90 preserved synthetic not-run states, and aggregate score/pass rejection. No actual artifact, winner, human outcome, or efficacy result is claimed.
- 2026-08-15 — Six decisions made answerable; none inferred. Added one fail-closed owner/provider register with 17 admissible paths, four authority roles, and a separate final release. Kept all selections, release, spend, calls, contacts, and efficacy explicitly unresolved or zero.
- 2026-08-15 — Authority typed; no human result inferred. Added five
fixture-specific role maps, six nonmechanical outcome lanes, 30 explicit
mechanical-pilot
not-runstates, and fail-closed human-authority checks. Kept scope selection, named people, model calls, and every human verdict unrun; the preflight remains at six blockers. - 2026-08-15 — One mechanism prepared; no layer selected. Decomposed two public skill bundles into one post-artifact evidence-reconciliation candidate with one insertion, one failure target, five evaluator-only mappings, and explicit exclusions. Corrected its schema so an unresolved or discrepant checkpoint cannot report overall pass. Kept owner ratification, model calls, and efficacy unrun; the preflight remains at six blockers.
- 2026-08-15 — Neutral runner isolates task input from evaluator truth. Added five deterministic source checks, 15 isolated fixture-arm envelopes, explicit model/evaluator separation, external output hashing, and fail-closed receipt identity validation; reduced the preflight from seven to six blockers while keeping zero calls and no spend authority explicit.
- 2026-08-15 — Current-harness preflight blocks premature comparison. Added exact packet/tool/skill custody, a common receipt schema, and a seven- blocker preflight; recorded zero calls and no spend, and made the neutral artifact runner the next executable gate.
- 2026-08-15 — Interactive packet completes the five-of-five source gate. Added two synthetic data versions, canonical state, ten expected states, seven replay trajectories, stale-selection recovery, exact export, history, keyboard, narrow delivery, and prepared-change contracts; kept application, accessibility, maintenance, human-outcome, and architecture claims unrun.
- 2026-08-15 — Specialized uncertainty packet makes four of five contexts executable. Added exact official ACS estimate, margin-of-error, geography, comparison, refusal, branch, render, and expert-review receipts; kept the baseline, expert verdict, reader outcome, and cross-domain claim explicitly unrun.
- 2026-08-15 — Governed-BI packet makes three of five contexts executable. Added a first-party synthetic semantic model, vintage, roles, verified queries, refusals, and exact fanout incident; kept production security, reliability, adoption, and human outcomes explicitly unrun.
- 2026-08-15 — Exploratory packet makes two of five contexts executable. Added the exact CC0 source, pooled/within-species reversal, branch-preserving task, and familiar-data use limit; architecture comparison remains gated on the then-missing packets and equivalent receipts across every arm.
- 2026-08-15 — First E0 source packet. Added the reconciled historical custody boundary and kept architecture comparison gated on missing source packets plus equivalent receipts across every arm.
- 2026-08-15 — Repair is a testable rung, not maintenance proof. Added the current DV-World repair and evolution results and kept their production, people, and cost boundary explicit.
- 2026-08-15 — One same-artifact afterlife, not a complete lifecycle. Added the KubeStellar production incident, repair, restored-deploy, and prevention chain plus PM4Py-UCM’s measured refinement-cost fragment; kept comparison, whole cost, handoff, accessible readers, decisions, and trust open.
- 2026-08-14 — Readability and navigation update. Added selective emphasis to the load-bearing conclusions and linked claims to the relevant visual experiences and deep-dive reports.
- 2026-08-14 — Initial public edition. Condensed the research, practitioner, human-skills, learning, agent-skill, specialized-vision, and forecasting work into one decision-oriented summary.