Explain a bounded account of the evidence.
Helpful: exploration, annotation alternatives, accessible implementation, source checks.
Danger: fluent generation substituting for reporting, authorial judgment, or reader testing.
Context
Preserve the question, data, decisions, artifact, provenance, and acceptance evidence. Then match the authoring and evaluation method to the human purpose.
Helpful: exploration, annotation alternatives, accessible implementation, source checks.
Danger: fluent generation substituting for reporting, authorial judgment, or reader testing.
Helpful: governed measures, stable layout edits, anomaly explanation tied to source visuals.
Danger: silent changes to metrics, thresholds, state, or hierarchy.
Helpful: semantic grounding, visible intermediate data, reusable verification.
Danger: confident answers over weak metadata, ambiguous measures, or hidden filters.
Helpful: many cheap views, branches, undo, direct manipulation.
Danger: turning the first plausible pattern into the final narrative.
Helpful: restricted representations, executable transformations, linked views, domain checks.
Danger: visual plausibility standing in for scientific correctness.
Helpful: adjustable assumptions, guided interaction, feedback, multiple representations.
Danger: generating interactivity without measuring what learners understand.
The lifecycle test
Question → data authority → generation → checking → correction → delivery → reader use → sustainment. The spine stays fixed; acceptance changes with the environment.
Current diaries expose source work, decomposition, defects, correction, mobile checks, publishing, and proposed updates.
Independent verification, editorial review, reader comprehension, accessibility, later refreshes, and civic use.
Source-to-claim trace, editorial gate, delivered desktop/mobile/access states, audience test, correction policy, and update owner.
Product contracts expose semantic models, queries, permissions, and review surfaces; testimony explains the need for narrow, owned metrics.
Routine task success, review queues, refusal, one-number incidents, trust recovery, and total cost in one independent study.
Approved metric contract, exact query and filters, permissions, owner, refusal behavior, escalation, and decision outcome.
Three same-artifact cases cross public delivery into a later event; two reach accepted repair, corrected delivery, and maintainer recheck.
Affected-user or independent-operator recovery is 0/3, transferred authority is 0/3, and complete twelve-state rows are 0/3. Whole cost, accessible-reader use, decision outcome, and calibrated trust also remain missing.
Keep provenance, release, exposure, event, diagnosis, repair, corrected delivery, maintainer recheck, affected-actor recheck, prevention, contribution, and authority separate.
One 117-person experiment measures immediate comprehension; one classroom case shows direct instruction rescuing failed AI-assisted construction.
Delayed construction, unfamiliar transfer, ordinary assistant use, and a complete learner-to-reader project.
Immediate and delayed unassisted performance, diagnosis, transfer, confidence calibration, and the eventual audience’s comprehension.
Controlled studies measure low-vision smartphone and blind nonvisual outcomes. MAIDR adds pre-AI and abstract-level AI study evidence plus a version floor. Graphy adds the same three blind co-designers returning for 12 sessions across four workshops and eight months while the interaction changed.
MAIDR's tested build and model remain unknown. Graphy is not a formal usability or performance evaluation; its workshops have no immutable code/model binding, and its seven-commit repository has no tag, release, or representative recheck.
Keep co-design, formal evaluation, study build, first containing release, changed release, recheck, and rewrite separate; compare exact AI versions with representative users on their own assistive technology.
Put the evidence to work
Three same-artifact cases join an AI-assisted dashboard to a later event, and two reach maintainer-verified restoration. None receives an affected-actor recovery receipt or transfers maintenance authority. No captured episode also joins whole cost, accessible reader use, a consequential decision, and calibrated trust.
Preserve AI provenance, pre-event delivery, actor-separated exposure, later event, diagnosis, accepted repair, corrected delivery, maintainer recheck, affected-user or independent-operator recheck, prevention, accepted second-person change, and transferred authority. Zero case clears all twelve. OpenClaw issue 30 and PR 31 remain open; Prism issue 81 is closed after a maintainer check but without the affected operator's return. Ticket status is not a recovery receipt. Read the complete account →
The CHI 2024 MAIDR study gives 11 blind participants a pre-AI multimodal chart system; the current project declares a complete TypeScript rewrite and an AI description layer. That is a valuable baseline plus a new artifact state—not an exact-version participant recheck. A public-health copilot paper describes a 16-person trust/usability method, then says the experiment was removed and reports no human result. Join exact artifact and version, AI state, release, later event, actor, custody, and outcome before counting a lifecycle.
Public code narrows the later eight-participant AI study to a legacy box-plot surface and v2.10.0 as the first containing tag. It does not name the participant-tested build or model. Later model swaps, verification fixes, deprecation, and the separate TypeScript rewrite trigger new representative checks; they do not supply them.
Graphy follows the same three blind co-designers while selected interaction changes are implemented between rounds. The final own-data workshop shows the evolved interaction in use, but the paper explicitly is not a formal usability or performance evaluation. Its seven-commit public repository has no workshop-version binding, tag, release, or representative recheck. Count repeated co-design—not a release afterlife.
Every required receipt appears somewhere in the held evidence, but none stays attached to one artifact lineage. A complete row needs repeated representative use, immutable tested build, exact exposure model, versioned release, later material event, and representative or actor-separated post-change recheck. This is zero of seven in named surfaces through 15 August 2026—not a claim about every private or inaccessible case.
Three returning blind co-designers across 12 sessions and eight months.
Workshop build and exact model; no release.
Bind the workshop, then recheck a changed release.
BLV study, study surface, version floor, and later maintenance.
Participant-tested build and model.
Return-user check after a pinned maintenance change.
Repeated repair attempts, v1.8.11, maintainer runtime check.
Exact model and affected-operator final recheck.
Operator exercises the corrected release.
Tagged release, non-owner failure, open repair proposal.
Corrected release and reporter recheck.
Merge → release → reporter reconciliation.
Live regression, repair, restored production, prevention.
Exact model and actor-separated recheck.
Independent user or operator recovery check.
Frozen v0.7.4 ledger: 18 sessions, 151 commits, 20 releases.
Representative use and elapsed field event.
Recipient returns after a released change.
50 native repairs and 80 changed visualizations.
Delivered artifact and returning user.
Carry one accepted output into real afterlife.
A 2021–2025 HealthTech program followed 21 projects through 84 recorded calls, notes and decision logs, backlogs and issue trackers, one governance-dashboard case, and a 16-startup survey. The case reached deployment with patients. It is longitudinal visualization evidence—not a generative-authoring study—and it does not measure total cost, accessible reader outcomes, calibrated trust, or comparative maintenance efficacy.
DV-World adds 50 native Excel repair tasks and 80 new-data or evolving-requirement tasks. The best reported agents reached 48% repair success and 51.44% evolution. Use a prepared change as a pre-delivery gate; the benchmark does not follow an accepted artifact through elapsed maintenance, handoff, reader use, or whole cost.
A four-month KubeStellar Console report and commit-pinned QA record follow one AI-assisted dashboard through a live blank-page regression, two-step repair, restored deployment, and a manually added post-build check. This is project-authored single-maintainer evidence—not a comparative rate, whole-cost ledger, handoff, accessible-reader result, consequential decision, or trust result.
OpenClaw Agent Dashboard says it was built with Claude Code and shipped v3.0.0. Thirteen days later, a non-owner operator reported that its cost view showed $0 on a custom provider despite present token data. The owner acknowledged that configuration had not been tested. A different non-owner's open PR 31 proposes a matching fallback, but has no maintainer review, merge, corrected release, or reporter recheck. A proposal is not recovery.
Prism says its family dashboard was built with Claude Code under human product direction. One non-owner operator's Home Assistant route failed at install, repair build, and later startup before they used ordinary Docker instead. v1.8.11 and its entrypoint contain the named socket and schema-replay repairs; the maintainer reports install, start, and restart checks on real Home Assistant OS. The operator never rechecked it. Route substitution is not recovery of the failed route, and no defect is attributed to Claude Code.
Prism's non-owner PR 23 carries a Claude-assisted weather visualization through owner-found defects, contributor repair, owner integration, and a credited release; the contributor later returns with accepted PR 41. OpenClaw's non-owner PR 15 puts Claude-assisted multi-provider pricing into the product after owner integration. These establish second-person change—and one repeat contributor—not release, incident, or recovery authority.
Non-owner changes the exact artifact; owner accepts it.
Prism PRs 23/41; OpenClaw PR 15.
Ownership or incident duty.
Same non-owner returns after elapsed time.
Prism, once after 17 days.
Durable team membership.
Named second person can release or respond and exercises that role.
Missing.
A contributor is a maintainer.
Affected or separate operator exercises the corrected route.
Missing in Prism issue 81 and OpenClaw issue 30.
Open patch or owner check is recovery.
Project-authored live failure, two-step repair, restored deploy, prevention check.
Independent acceptance, whole cost, handoff, accessible readers, decisions, trust.
Keep fix, deploy, live recovery, and future gate as separate receipts.
Non-owner report plus a different non-owner's open repair proposal.
Maintainer acceptance, merge, corrected release, reporter recheck, invoice reconciliation, downstream outcome.
Proposal is not accepted repair; close only after release and independent recheck.
Non-owner install, repair-build, and startup failures; corrected release; maintainer runtime check.
Affected-operator recheck, whole cost, handoff, accessible readers, decisions, trust.
Keep report → repair → release → maintainer recheck → independent recheck as five receipts.
Intended audience, decision, source authority, local definitions, baseline.
Can the creator explain the claim? Who owns the metric, judgment, and consequences?
Were task, stakes, expertise, and current comparison declared before use?
Prompts, transformations, direct edits, failures, waits, rollback, rejected output.
What was inspected or repaired? What review changed the artifact, and what stayed disputed?
Measure time to first candidate separately from time to accepted work.
Named approver, acceptance contract, published state, desktop, mobile, keyboard, and assistive-technology checks.
Did the real surface pass? Keep approval, publication, and reader use as separate states.
Recompute claims and capture the delivered state independently.
Defined reader tasks, comprehension, decisions, confidence, errors, feedback.
What did intended readers understand and do? Who was excluded?
Measure decision quality, calibration, recovery, and accessibility—not satisfaction alone.
Refresh host, dependencies, second maintainer, incidents, correction policy, retirement, human and model cost.
Can someone else update or retire it? Who owns the next refresh and failure response?
Follow a dependency change or handoff; compare whole cost against current direct work.