Reader experience

The reader sees a claim—not the prompts, repairs, or uncertainty behind it.

Creator delight is not a proxy for reader comprehension. The visual artifact inherits expectations from journalism, science, business, education, or social media.

Perceived trust

Clarity and familiarity matter—but neither proves truth.

In a 2025 study, 31 of 37 participants mentioned clarity while ranking charts. Sources, integrity, familiar forms, and aesthetics also mattered, with substantial differences between people. The study measured deliberative trust, not correctness, comprehension, or behavior.

Trustworthy by Design ↗
Provenance labels

“AI-generated” is not a universal trust switch.

In a preregistered 2021 experiment, people brought strong preferences for human or algorithmic recommendations, but data relevance usually dominated actual selections. The same label suggested precision to some and missing human judgment to others.

Vis Ex Machina ↗
Model evaluation

A model that decodes a chart is not a human-reader test.

On 60 synthetic charts, multimodal models enumerated data structure while 24 people formed trend narratives and reacted more to layout and overlap. Each could succeed on a different notion of intent.

How Do LLMs See Charts? ↗
Evidence lineage

Name the AI role before naming the human outcome.

Across ten held rows, one directly carries selected AI-generated charts into an independent controlled-reader effect. Five measure adjacent human outcomes where AI labels, explains, assists, or sits beside a chart; three concern creators or co-designers; one uses VQA without people. Zero reaches accepted delivery and a later same-lineage reader recheck.

Read the eight-receipt boundary →
Reader scaffolding

Guiding attention can outperform merely answering.

In a randomized 117-person experiment, data stories, passive AI Q&A, and proactive scaffolded dialogue all improved visual-comprehension scores. After support was removed, the proactive group’s median was 6/6 versus 5/6 for both alternatives; completion time did not differ.

GenAI agents and visual comprehension ↗
Measured harm

A data-consistent chart can still induce wrong answers.

In a 48-person controlled experiment, second-phase accuracy was 71.9% for readers shown AI-generated misleading charts and 88.3% for readers shown correct charts. This measures bounded chart-QA harm—not prevalence, calibrated trust, or consequential decisions.

ChartAttack ↗
Accessible delivery

The device, modality, chart type, and interaction all change the outcome.

A 26-person low-vision smartphone study found 100% completion with a full interactive treatment versus 61.5% with a screen magnifier. A separate 10-person blind-reader study found similar device-level accuracy but different time, workload, and chart-type results. Neither tested AI-generated charts.

Longitudinal evaluation boundary

15 named cases · 1 development-decision study · 1 operational-pipeline near-miss · 0 full episodes.

Lexara follows six CVA developers using a deployed evaluation toolkit for two weeks across 38 experiments, 57 newly authored cases, ten models, and six prompts. They made and challenged model/prompt selections using their own data, but the study does not join an immutable build and exact configuration to accepted downstream delivery, an intended-reader decision or calibrated trust, whole cost, and a later post-release recheck. Keep development selection, contract acceptance, audience consequence, confidence calibration, and return use separate. Read the twelve-receipt account →

Artifact custody boundary

One pinned output bundle · zero audience delivery/rechecks.

One explicit GPT study among 122 stable-key titles adds a pin-able, 1,271-blob supplement of prompts, results, generated code, and grader bundles. It does not add an immutable provider snapshot, versioned release, accepted final project, delivery to the government, UN, or student audiences named in prompts, an audience outcome, or a later same-lineage return. Treat the other 121 title nonmatches as not surfaced, not excluded. Read the complete custody boundary →

Non-AI afterlife boundary

Ten years of public delivery · zero transferable AI receipts.

AIDSVu adds aggregate use, named planning applications, recurring governance, explicit public-data states, and a later 2026 data release. Those are mature platform receipts. The held surfaces expose no AI visualization role, immutable build, version-bound measured audience outcome, affected-audience recheck, or whole cost. Keep it as a non-AI comparator and keep 86 abstract cue nonmatches unexcluded. Read the complete afterlife boundary →

Primary human-evidence boundary

Four empirical cases · three receipt lanes · zero deployments.

Five recovered full texts separate controlled audience measurement, public-sector co-design/demo/bounded use, and enterprise demonstrator plus production intent. None binds an accepted field release to a later affected-audience recheck, and all five full-text GenAI cue screens are empty. Later recovery adds 20 substantive surfaces, including one further full chapter, and leaves two DOI rows content-unassessed. Co-design, a demo, a human task study, and production intent are different receipts—not deployment. Read the complete receipt ladder →

Field-use boundary

One non-AI operational comparator · zero complete AI lifecycles.

Eight stable non-DOI keys yield three DOI repairs, two year corrections, one rejected foreign PMID, and four full texts. The EMR cancer diary reports increasing system-log use; an independent review recovers an 11-clinician QUIS median of 4.38. That is meaningful embedded-use and usability evidence, but it supplies no AI role, immutable accepted build, patient or decision outcome, later event, or affected-clinician recheck. One primary text remains gated and unassessed. A field-use receipt is still not an AI lifecycle. Read the complete field-use boundary →

Human-evaluation boundary

Seven human evaluations · zero complete production lifecycles.

Nine high-signal DOI rows yield eight substantive primary abstract or official-project surfaces and seven participant/evaluator studies. InfoViP is the delivery near-miss: seven FDA safety evaluators shaped and evaluated the prototype, suggestions were addressed, and an official page says an enhanced NLP and unsupervised-learning version will be installed in production. Future tense is not installation, acceptance, routine use, regulatory outcome, later change, or evaluator return. Subsequent recovery reduces the content-unassessed DOI remainder from 14 to two. Read the complete human-evaluation boundary →

Production component + authority boundary

Approved component · 3 research mentors · 0 operating authorities · closed ≠ shipped · 0 full.

The final 2025 CIOMS report supports an approved AWS/AERS component processing more than 30 million historical plus about 8,000 daily submissions while leaving a solid QA plan, completed audits, routine roles, and signal effects open. The now-closed Elsa/API/UI opportunity names Joshua Xu, Leihong Wu, and Oanh Dang as research mentors, but supplies no selection, work, acceptance, or release receipt. A fellow is a nonemployee barred from inherently governmental functions. Neither the actor profiles nor an adjacent AI-QA project assigns InfoViP operation, maintenance, QA execution, release approval, authorization, a completed audit, or the Elsa application join. Read the actor-and-authority boundary →

Residual primary-content boundary

14 investigated · 11 substantive · 1 full chapter · 3 metadata-only at that pass.

A full chapter connects three UX experts and 25 distinct problems to an implemented third version. A corridor stakeholder case and university-network case add practice context. None exposes a GenAI visualization role, accepted field release, routine-use denominator, consequential outcome, or later affected-actor return. Expert-evaluated, redesigned, and applied with stakeholders are useful stages—not synonyms for deployment. Read the complete residual account →

Public-delivery afterlife boundary

3 rechecked · 1 newly substantive · 1 live interface · 2 unassessed.

B110 now has an official abstract, pinned framework source, a pinned DiscoverWater application, and a same-named KU interface reachable in August 2026. At that pass, B92 and B91 remained content-unassessed. The live bytes are not bound to a SHA; acceptance, continuous or ordinary use, audience outcome, recheck, whole cost, and an AI role remain absent from held surfaces. Read the complete delivery boundary →

Live-build lineage boundary

1 live page · 0 exact page builds · 2/3 sampled assets canonically equal.

The dated live page differs from the sole published v1.2 page state; the linked v2.0 implementation is R/Shiny. Eight of twelve dependency names occur in pinned application source, and two of three sampled data assets match canonically. The partial lineage is real, but no manifest binds the deployed page to source. A prototype demonstration plus analytics and comment hooks add no acceptance, use, audience outcome, accessibility, later-return, whole-cost, or AI receipt. Read the complete source-lineage boundary →

Primary recovery + stop rule

One abstract assessed · one row access-blocked and unassessed · generic search paused.

B92's exact publisher abstract describes multicriteria risk evaluation, Monte Carlo simulation, Kendall's tau rank comparison, and graphs for ordering pipeline sections. B91 remains content-unassessed after three bounded lawful passes over six named surfaces. The content ledger stays 6 full / 20 abstract or official / 1 unassessed / 0 complete AI lifecycles; the separate workflow state is zero active generic B91 targets. Reopen only when exact new lawful custody appears. Paused is not negative. Read the complete stop-and-reopen boundary →

Review lineage boundary

Full review: 122 keyed · 5 authority gaps.

The complete register binds 122 of 127 supplement positions to distinct A208 review keys. Five exact-looking DOI identities have no A208-controlled join, and one admitted key carries a PubMed identifier that points to a different paper. Keep source label, review key, candidate, identifier validation, and admission authority separate. Lifecycle screening is unstarted. Never promote the five tempting matches, reuse the conflicted PMID, score an unresolved row, or report zero of 127. Read the complete boundary →