Research source document · Evidence reviewed through August 18, 2026

The state of AI-assisted data visualization research

Status: research snapshot, evidence cut 2026-08-18. Recheck by 2026-11-16, or earlier after a material model, harness, benchmark, registry, or analytics- assistant release.

Short version: executive summary. Forward view: What would unlock the next capabilities in AI-assisted data visualization? turns the moving-baseline evidence into dated, scoreable forecasts and keeps model-driven and technique-driven gains separate. Assessment view: How close is AI-assisted data visualization to the ideal? defines a context-sensitive ideal, scores nine current capabilities on one fixed maturity scale, reconstructs the same scorecard from 2017 through August 2026, and translates the forecasts into explicit score movements.

This report is for researchers, product teams, designers, and engineers who need to understand what AI-assisted visualization systems can currently do, which techniques appear to improve them, and how those claims are measured. It is a research synthesis and mechanism comparison, not a product ranking or buyer’s guide. Tools, agent skills, and commercial BI assistants appear as examples of techniques in practice; their documentation does not establish comparative quality.

Eight open-source repositories were inspected statically at the pinned revisions listed below. Eight core system papers were read in full, four initial benchmark and training studies were read for the moving-baseline analysis, three current commercial product surfaces were checked against first-party documentation, seven first-party model and harness release records were used for the shared timeline, and earlier papers in the underlying research packet were re-read where they controlled a comparison. Three established visualization frameworks were added to ground the distinction between task goals, narrative explanation, dashboards, and exploratory tools. A companion deep reading adds 18 agent-skill files or families, broader skills benchmarks, and the paired SciVisAgentSkills study; its report keeps registry installations, repository popularity, and behavioral evidence separate. A second specialized-vision report separates chart QA, parsing, OCR/layout, grounding, quality or integrity critique, and verifier/repair roles rather than treating every visual model as an interchangeable judge. A third companion, Learning data visualization when AI can make the chart, separates assisted performance from retained skill, distinguishes productive difficulty from removable friction, and forecasts how human practice should be reallocated as implementation becomes cheaper. A fourth companion, What AI-assisted data visualization benchmarks actually measure, translates benchmark names into their inputs, outputs, source distributions, graders, human context, supported claims, and nonclaims. No external repository, package, installer, skill, model gateway, or untrusted script was run. Repository tests and included evaluation results are observations about what the source contains, not independently reproduced results.

Start here: what agentic data visualization is trying to do

Agentic data visualization is not one task. A system may help a person discover a pattern, explain a finding, monitor a changing situation, make a decision, or produce and revise a durable visual artifact. These purposes can occur in a sequence—exploration may eventually become a published explanation—but they do not have the same success condition.

Rather than invent a new top-level taxonomy, this report borrows the discipline of Brehmer and Munzner’s established visualization-task typology: describe why the work is undertaken, what data and outputs it acts on, and how encoding and interaction support it. This matters acutely for agents. Returning a technically valid chart answers part of “how”; it does not show that the system served the intended human purpose.

Primary goal Typical environment What good support looks like What the agent must not optimize away
Explain and present journalism, public explanation, narrative graphics evidence supports a specific account; annotation, sequence, surrounding prose, and visual emphasis work together; defined readers can follow it authorial claim and context, counter-reading, integrity, accessibility, publication judgment
Explore and discover analyst work, open-ended research, exploratory tools transformed data and alternatives remain visible; the person can steer, branch, correct, and reverse course ambiguity, provenance, direct control, and the ability to abandon a bad line of inquiry
Monitor and respond operational dashboards, shared awareness surfaces current state, thresholds, anomalies, and action state are legible and trustworthy freshness, stable measures, permissions, escalation rules, and operational consequence
Compare and decide strategic and analytical BI, decision support authoritative measures, fair comparisons, alternatives, uncertainty, and tradeoffs support a consequential judgment semantic custody, metric ownership, role access, and human accountability
Produce, revise, and reuse authoring tools, coding agents, scientific or interactive applications the artifact executes, remains editable, reproduces its computation, and behaves correctly in delivery domain constraints, edit locality, browser state, specialized scientific meaning, and final acceptance

These are overlapping goals, not product bins. Segel and Heer characterize narrative visualization by the balance between author-driven narrative flow and reader-driven discovery. Sarikaya and colleagues show that dashboards also vary by purpose, audience, interaction, and data semantics, and differ materially from exploratory visualization tools. A newsroom interactive may contain both guided explanation and exploration; a dashboard may monitor, analyze, teach, or communicate. The important step is to declare the sequence and apply the right evidence at each transition.

That produces a general requirement: every system should preserve the question, data, decisions, artifact, and evidence. A context contract should then select the appropriate authoring and evaluation method. A universal generator, critic, or acceptance score would erase differences the system needs to understand.

What is being compared

This landscape contains several kinds of thing that are easy to blur together. A paper can describe a tool, a repository can implement a paper, and a benchmark can be used to test several tools. They still make different kinds of claim.

Kind of object Plain-language meaning Examples in this review What inspecting it can establish
Research paper or system A proposed way to generate or edit visualizations, usually paired with an experiment Raiven, NL2Dashboard, nvAgent, DashChat, PlotGen, NL4DV-LLM, Data Formulator 2 What method was proposed and how it performed under the paper’s tasks, models, baselines, and graders
Open-source tool or framework Runnable code intended for people or applications to use Data Formulator, Lumen, LIDA, Vizro What the current implementation can do and what checks it contains; not necessarily whether it improves human outcomes
Agent skill or plugin package Instructions, references, templates, and sometimes scripts that shape how a general coding agent works Vizro’s end-to-end flow, OpenAI visualize-data, AntV skills, Markdown Viewer Vega skill, claude-skillz What workflow and verification the package requests; only a with/without behavioral test can show whether an agent follows it or benefits
Benchmark or evaluation harness A fixed set of tasks plus execution and scoring rules DashArena, VisEval inside nvAgent, MatPlotBench inside PlotGen Comparative performance under that harness; not automatic proof of production reliability or reader benefit
Commercial BI assistant A vendor feature operating inside an established analytics product Power BI Copilot, Tableau Agent, Looker Conversational Analytics The current product contract and native sources of grounding; public documentation is not a comparative accuracy study
Technique A reusable mechanism that may appear in any of the objects above typed intent, compilers, semantic grounding, browser replay, direct manipulation A candidate design principle. Its value must still be tested in the environment where it will be used.

The objects fit together roughly like this, but the short pipeline hides who may decide what:

question + data
      |
      v
authoring system or product  <--- an agent skill can guide this work
      |
      v
structured plan, DSL, or code
      |
      v
renderer + browser  ---> executed chart or interactive dashboard
      |
      +--- deterministic checks test values and behavior
      +--- a benchmark supplies repeatable tasks and scoring
      +--- a human study tests how people author, understand, or decide

A reference anatomy assigns information and authority

A bounded audit of six held primary authoring or analysis systems—Data Formulator 2, Raiven, MatPlotAgent, Text2Vis, ViviDoc, and interactive task decomposition—supports an eleven-stage reference anatomy. The stages are a custody checklist, not one required agent topology:

Stage Information and legitimate authority Boundary
Intent A person declares purpose, audience, requested artifact, stakes, and stopping condition. A prompt is not source, editorial, or publication authority.
Data and semantic grounding Owners and deterministic contracts bind named sources, fields, values or metadata, units, measures, permissions, and uncertainty. Fluent field selection cannot redefine an authoritative measure.
Inspectable representation Assumptions, transformations, plan, state, encodings, interactions, constraints, and history can be inspected and rejected. Hidden reasoning is not a governed specification.
Generation A declared model and context propose code, specification, content, or an answer. A candidate is not execution, correctness, or permission to publish.
Execution A compiler, runtime, or browser turns exact code or a DSL into observable behavior and pixels. A successful run does not prove semantic fidelity or reader value.
Critique Deterministic checks, specialist models, and humans test named and different failure classes. A critic is a bounded sensor, not a universal quality or acceptance authority.
Repair and history Defects, feedback, changed state, prior state, non-regression results, and cost remain traceable. A proposed repair is not a verified fix.
Human control A named person can inspect, edit, branch, stop, reject, or escalate at consequential boundaries. The opportunity to intervene is not a recorded acceptance decision.
Accepted delivery An exact build, frozen rule, named approver, deployment, and delivery receipt travel together. Benchmark pass, satisfaction, export, and local render do not qualify.
Reader or decision outcome Intended people encounter the delivered surface under declared device, access, and task conditions. Author ratings and model judges cannot stand in for readers.
Later maintenance A later data, model, requirement, browser, or ownership event is repaired and rechecked with authority and whole cost retained. Initial success does not establish updateability or handoff.

The six-system audit finds 6/6 generation, execution, and some critique route; 4/6 meaningful human-control gates; and 0/6 exact-configuration accepted delivery, intended-reader outcome, or later maintenance. The 6/6 critique count is not architectural convergence: typed validation, runtime errors, visual-model feedback, and human inspection see different evidence and hold different authority. VisJudge can add one specialist critique signal; DashboardQA can test downstream agent interaction. Neither closes the human gates.

Primary anchors: Data Formulator 2, Raiven, MatPlotAgent, Text2Vis, ViviDoc, and interactive task decomposition.

Public acquisition is not adoption

The same custody rule applies to agent-skill popularity. Twelve named data-visualization registry listings carry 37,720 listing-summed install signals at the 14 August evidence cut. They collapse to ten parent repositories and eleven documented provenance lineages. The sum is not a count of unique people, teams, successful installations, invocations, or outcomes.

The registry calls its field a total deduplicated install count but does not publish the deduplication key or interval. At a pinned official CLI commit, an install event is assembled from selected skills after target results are collected; telemetry may be disabled or absent, and its schema contains no skill-invocation, task-outcome, retention, or organizational-acceptance event. The private audit therefore finds 12/12 public install signals and 0/12 held rows with observed real-task invocation, repeat use, organizational acceptance, or outcome/afterlife. Those zeros are bounded to the named public surfaces; private use remains unknown.

Stage Receipt needed Named public rows
Listed stable listing and captured package 12/12
Install signal dated registry definition and count 12/12
Successful presence exact package and version present after installation 0/12 observed
Real-task invocation task, model, harness, package, and execution receipt 0/12 observed
Repeat use or retention same actor or environment returns, with a denominator 0/12 observed
Organizational acceptance named owner, governed scope, and approval 0/12 observed
Outcome and afterlife accepted work, comparison, cost, later event, and recheck 0/12 observed

Use the skill-package report for the package-level table and evidence boundary. Use registry counts to choose what to inspect, not to make an adoption, procurement, or effectiveness claim.

Available controls are not governed deployment

The organizational-governance evidence also separates into two rails. Power BI/Fabric, Tableau, and Looker document meaningful enablement, data-boundary, and monitoring controls. Provider stories separately report named-feature use of Looker Conversational Analytics at Google Cloud Support and Copilot in Power BI among KPMG developers. Two longitudinal studies add adjacent governance- process evidence. Across those seven held rows, zero joins the complete same-deployment record.

Evidence rail Held rows What it establishes What it does not establish
Provider control contract 3 controls an organization can configure the organization’s effective configuration, approval, use, or outcome
Named-feature organizational use 2 attributable provider-published reports of routine use exact controls, audit and incident disposition, independent acceptance, or later recheck
Adjacent governance process 2 longitudinal authorization, evidence, ownership, or audit practice around organizational AI an AI-assisted visualization deployment using the named provider feature
Complete governed deployment 0 — no held row clears all nine receipts

The join requires one feature and version, organizational authorization, enablement and scope, data or semantic authority, monitoring and audit, routine use, incident or exception disposition, an independently accepted outcome, and a later recheck. Provider controls from one row cannot be spliced to outcomes from another. This is a bounded public-surface result, not evidence that no organization holds a stronger private record. The practitioner account contains the provider-control, organizational-use, and source boundaries.

The eight core papers are not eight versions of the same product. Additional benchmark and training studies are introduced later where they answer the separate question of how the model baseline is changing.

Paper What it is, in plain language Why it is in this comparison
Data Formulator 2 A study of a chart-authoring interface where a person directly chooses visual encodings and asks the model for missing data transformations. It is the clearest small study of mixed-initiative authoring, visible intermediate data, branching, and verification behavior.
Raiven A scientific-visualization prototype where the model writes a restricted visualization language and a compiler produces linked 2D, 3D, and table views. It supplies unusually strong bounded evidence that a DSL and compiler can eliminate many code-generation failures when the desired view is already specified.
NL2Dashboard A dashboard architecture where the model produces a compact structured plan and uses small edit operations instead of rewriting the whole application. It tests whether an intermediate representation improves controllability, edit locality, and token use.
nvAgent A natural-language-to-database-visualization pipeline that prepares a schema, composes a visualization query, executes it, and repairs failures. It provides a large benchmark and useful ablations for structured composition and execution-guided validation, especially across multiple tables.
DashChat A conversational industrial-dashboard mockup tool with a dashboard DSL, design-pattern retrieval, structured edits, and history. It combines a held-out prompt evaluation with a small user study of rapid prototyping and human correction.
NL4DV-LLM A method that turns a natural-language question into an inspectable analytic specification: selected fields, analytical task, and candidate visualizations. It shows why preserving explicit mappings and multiple interpretations can be more useful than jumping directly from prose to chart code.
PlotGen A scientific plotting pipeline with separate numeric, text, and visual feedback passes around generated Matplotlib code. It is evidence for multimodal feedback, while also illustrating why extra agents and extra model calls need an equal-budget causal test.
DashArena A benchmark, browser executor, and human-calibrated judge for interactive dashboards—not a dashboard-authoring product. It adds a serious measure of task-grounded analytical and interaction quality by replaying the author’s intended interaction sequence.

Terms used throughout the report:

State of the field

This synthesis distinguishes repeated findings from environment-specific choices, direct negative or null evidence from merely unsupported claims, and demonstrated capability from open research questions.

Findings in one page

The field is not converging on a single agent architecture. It is converging on a more useful engineering pattern: reduce the part the model must improvise, externalize its decisions, and test the resulting artifact at the layer where failure can occur.

The strongest current techniques are:

  1. A typed visualization or dashboard representation between intent and rendering. Raiven, NL2Dashboard, ViviDoc, nvAgent, NL4DV-LLM, current Data Formulator/Flint, DashChat, and the best library-specific skills all constrain generation through a DSL, intermediate representation, or declarative spec. This improves syntax, edit locality, token efficiency, and the ability to validate individual decisions. It does not establish that the underlying question or takeaway is worthwhile.
  2. Deterministic work for deterministic claims. Compilers, schema checks, aggregation checks, executable code, browser traversal, replayed interactions, and regression comparisons outperform asking a model to pronounce an artifact correct. The model remains useful for ambiguity, semantic mapping, critique, and repair; it should not substitute for arithmetic or runtime evidence.
  3. Grounding in the actual data contract. Governed BI products increasingly bind natural language to semantic models, field descriptions, verified queries, permissions, sample values, and business glossaries. This is more consequential than assigning a generic agent an analyst persona.
  4. Mixed-initiative authoring. Data Formulator 2 and DashChat make natural language one control surface among direct manipulation, structured selection, visible transformed tables, history, branching, and reversion. This has stronger human-use evidence than prompt-only generation.
  5. Evaluation over the rendered artifact and intended use. DashArena’s new contribution is not another chart score. A system authors an interaction trajectory; a browser replays it; the judge sees task, screenshots, schema, and execution evidence. Human agreement improves materially when interaction evidence is present. This is the closest current answer to the earlier gap around communicative and analytical value, but it is not a measure of reader comprehension, learning, retention, or decision quality.

“Multi-agent” is therefore not a transferable technique by itself. It earns a place only when a role has a different information boundary or tool contract: schema access, transformation execution, visual inspection, browser interaction, or independent acceptance evidence. A planner, composer, and validator can help because they manipulate different representations and evidence, not because their labels simulate a team. Equal-budget single-agent results and several ablation findings still argue against agent count as a default quality lever.

Where the literature converges

The convergence is a common engineering posture, not a standard product architecture:

Where approaches legitimately diverge

Several competing approaches can each be right in different environments:

Direct negative, null, and conditional findings

Unsupported shortcuts are a different category

An unsupported claim is not proof of harm. It is simply insufficient evidence for acceptance. The AntV saved “success” results did not include browser render outcomes. The prompt-only packages in the initial sample have no behavioral with/without proof. LIDA’s code-and-text evaluator does not inspect the rendered chart. A validator that fails open is not an acceptance mechanism. Stars, installs, code returned, a nonblank render, synthetic personas, or one vision-model score may be useful inputs or baselines; none establishes data fidelity, integrity, reader understanding, or decision quality.

The model and harness baseline is moving

Shared timeline: September 2023 through August 2026

The three lanes below align general model and harness releases with the changing scope of visualization systems and benchmarks. They are a shared clock, not a causal model: release dates do not prove that a platform milestone caused a paper’s result, and publication dates lag the work. Scores remain comparable only inside the named studies.

The web presentation places every milestone on one proportional September 2023–August 2026 axis. Marker position encodes the date; label width does not encode duration. On narrow screens, the same evidence becomes a vertical chronology inside each explicitly named lane.

Period General model and harness baseline Visualization systems and techniques Benchmark and evaluation frontier
Sep 2023 GPT-4V makes image input broadly available, enabling general models to inspect visual artifacts. The main public capability question is still dominated by producing or reproducing individual static charts. Existing evaluation mostly stops at code, structure, execution, or static-image similarity.
May-Oct 2024 GPT-4o adds native multimodality; Anthropic’s computer-use beta adds screen, cursor, click, and typing actions. Data Formulator 2 and NL4DV-LLM make transformed data, direct controls, history, and analytic specifications visible. Plot2Code tests plot-to-code reproduction; VisEval tests 2,524 natural-language visualization queries across 146 databases with heterogeneous checks.
Feb-Aug 2025 Claude Code brings a terminal coding agent; OpenAI’s agent tools bundle web, file, computer use, and tracing; GPT-5 targets coding and agentic work. PlotGen, DashChat, and nvAgent add multimodal feedback, dashboard-specific languages, schema planning, execution, and repair. Text2Vis combines data, questions, answers, code, and annotated charts across 1,985 tasks and isolates targeted feedback from generic prompting.
Jan-Aug 2026 Hosted shell, computer environments, persistent workspaces, and reusable skills become first-party platform primitives. NL2Dashboard tests a compact editable dashboard plan; Raiven tests a restricted scientific language and deterministic compiler. RealChart2Code adds real data and multi-turn refinement; CharTide compares newer general and specialized chart models; Dashboard2Code and DashArena add state and replay; Chartography and FinChart-Bench test difficult professional reading; ChartDiff and multi-chart PolyChartQA add cross-chart reasoning; Chart-MRAG adds chart-bearing documents and retrieval; multilingual POLYCHARTQA and MM-JudgeBench expose language gaps in readers and judges.

By the August 2026 evidence cut, the research question is no longer only whether a model can emit code for a plausible static chart. It is increasingly whether a model plus its harness can preserve meaning, construct a multi-view interactive artifact, exercise it in a browser, and supply evidence that it supports the analytical task. Reliability has not kept pace with that expanding scope.

The baseline has improved materially over the past 36 months, but the published record does not provide a clean visualization-specific learning curve from August 2023 to August 2026. Benchmarks changed, later papers often reran only a subset of models, judges changed, and current systems combine models with different prompts, tools, and test-time budgets. The defensible conclusion is a direction and a set of measured slices—not a universal annual improvement rate.

A fixed scorecard makes the remaining distance explicit

The companion capability scorecard applies one six-level rubric to nine dimensions: task framing; data, semantics, and provenance; construction; interpretation; integrity critique; steering and repair; interaction, responsiveness, and accessibility; reader or decision outcomes; and production, governance, and maintenance. The scale runs from “not demonstrated” through “outcome-proven.” It does not average the dimensions.

On that rubric, the August 2026 field is usable in bounded contexts for data grounding, ordinary construction, interpretation, and mixed-initiative repair; repeatable on bounded tests for integrity critique and delivered interaction; and only demonstrated—not established—for reader outcomes and the full production lifecycle. The same rubric can be applied retrospectively without pretending that unlike benchmark percentages share a common numerical axis. The historical reconstruction shows broad acceleration after 2023, but no common-core dimension yet crossing into field-wide deliverability. Context rows replace the generic ideal with the actual success condition for explanation, exploration, monitoring, decisions, production, scientific work, and reader assistance.

What the closest comparable slices show

The web presentation uses paired points for the before-and-after values below. Each measure keeps its own labeled scale and exact endpoints; readers may compare the two points inside a row, but not horizontal position across metrics. There is deliberately no combined score or common raw axis.

Plot2Code asks a multimodal model to reconstruct a scientific chart as code. An anonymous 2026 preprint re-reported direct, single-pass results on the same Python/Matplotlib subset and normalized scores over the full test set. Between the September 2023 GPT-4V result and the June 2025 Gemini 2.5 Pro result:

Measure Sep. 2023 Jun. 2025 Change over 21 months
Code executes 84.1% 87.9% +3.8 percentage points
Text in the recreated chart matches 48.5% 71.7% +23.2 points
Rendered-chart quality 5.45 / 10 7.65 / 10 +2.20

The large movement was in visual and textual fidelity, not basic code execution. This is a useful comparison, but it is one preprint’s reconstruction of one chart-to-code benchmark, not a field-wide time series.

CharTide, a peer-reviewed ACL 2026 training study, supplies a second same-paper comparison of GPT-4o and GPT-5 on three chart-to-code benchmarks. GPT-5 improved ChartMimic high-level similarity from 87.7 to 94.7, Plot2Code text match from 52.6 to 61.9, and ChartX’s five-point score from 2.61 to 3.59. Execution moved much less. The pattern is again uneven: generation is becoming more faithful, while harder visual reasoning and semantic alignment still leave substantial headroom. CharTide also shows that a small chart-specialized model can match or exceed a larger general model, so raw frontier scale is not the only route to improvement.

DashArena shows what the 2026 frontier can attempt that earlier chart benchmarks barely measured: a current general model can produce a multi-view interactive dashboard and a replayable intended-use trajectory from an open-ended task. Its top model was competitive with the anonymized human baseline in aggregate preference. Yet no model exceeded 86% render success or 74% replay success, and execution-clean semantic failures remained common. Breadth has expanded faster than reliability.

Dashboard2Code makes another part of dashboard behavior measurable. Its 180 Plotly Dash dashboard-code pairs cover 20 visualization types and eight callback patterns, with 450 interaction tasks. The best reported configuration scored 79.4 overall and 64.2 on the most complex interaction level. DOM access raised one model’s successful UI exploration from 20.61% to 64.58%, showing the value of exposing executable structure rather than relying on pixels alone. The benchmark also exposes a particularly dangerous failure: an interface can respond while hidden state or a transformation is factually wrong. It uses a fixed 1920×1080 viewport and excludes animation and popups, so mobile, responsive, and animated behavior remain outside the result.

Chartography asks a different question: can a model read difficult charts used in professional work? Its 100 practitioner-authored tasks span 12 domain labels and were independently verified by three experts per task. Thirty model configurations were run twenty times per task. The best tested configuration reached 45.0% mean pass@1. Greater reasoning effort helped in eleven of twelve paired comparisons, but the median gain was only 4.5 percentage points and longer reasoning often elaborated an initial visual misread. Because the tasks were deliberately screened for difficulty, 45% is not a prevalence estimate over all professional charts. It is strong evidence that success on basic visual-literacy or chart-QA sets does not imply dependable professional reading.

Scaffolding still helps, but generic instruction is already losing value

The most informative ablations do not say “more prompting is better.” They say that scaffolding helps when it adds a missing representation, tool, or evidence channel.

The current operational baseline is also no longer a bare model call. Coding harnesses increasingly arrive with planning, repository search, execution, browser inspection, and packaged visualization guidance. A no-skill model is still a useful diagnostic control, but it is not the realistic alternative to a application-specific extension. The decision baseline must be the current model plus the normal harness, its default tools, and any ambient skills or instructions.

A current evaluation-readiness audit shows why that distinction matters. The product-owned historical record now resolves 14 case definitions, 18 run directories, and 13 judge directories, but the current public evaluation surface has only the narrow editorial sentinel. Of the five contexts required for a shared baseline—editorial, exploratory, governed BI, scientific, and interactive—all five now have executable source packets. The exploratory packet uses the exact CC0 Palmer Penguins table to test a real pooled/within-species sign reversal, branch preservation, row accounting, and render receipts. Its widespread use makes it a known-failure sentinel, not held-out performance evidence. The governed-BI packet uses a first-party synthetic semantic model, four roles, six verified queries, four refusals, and a fanout incident that produces $16,830 against the authoritative $7,730. That makes permission and recovery receipts executable, not production security or reliability. The specialized packet uses official 2024 ACS B19013 estimates and margins of error. Iowa, Kansas, Montana, and Wyoming span only $192; their largest internal z-score is 0.149508 against a declared 1.645 threshold. An exact point-only chart can therefore be mechanically valid while its strict winner story is unsupported. The interactive packet uses two fictional Riverbend maintenance snapshots, five canonical state fields, ten exact states, and seven isolated trajectories. Its critical transition makes selected T007 ineligible and leaves only T012 visible; selection, detail, accessible summary, URL, and export must change together. It also contracts history, reset, keyboard, narrow delivery, and a prepared V2 change. Historical auditability and source-packet readiness are not baseline readiness. Do not compare architectures until the exact current model/harness manifest is frozen and bare, normal-harness, and one-added-layer arms can emit equivalent receipts.

The first manifest preflight locked the five packet hashes, local toolchain, available-skill hashes, task order, three arm roles, and one 18-field receipt schema. A credential-free runner now clears the mechanical runner gate: five source checks run before it stages 15 isolated fixture-arm attempts; model inputs are separated from evaluator-only truth; artifact, render, trace, and check files are hashed outside the arm; and receipt identities fail closed. The editorial case required an explicit redaction because its source file contains the answer key. A whole-packet copy would have changed the task.

The successor preflight does not call the comparison ready. Six conditions still block the first call: immutable served-model identity; exportable server- harness configuration; same-model bare access; an authorized capped budget; one versioned Vizier mechanism; and named acceptance authority or explicit not-run treatment. The runner and successor verifier record zero calls and no spend. A passing plumbing self-test, observed gpt-5.6 label, and local codex-cli 0.147.0 remain neither a model result nor a server-side reproduction receipt.

The same finding changes different decisions without changing the evidence:

Audience Decision layer
Practitioners and editors Show the estimate and interval; describe the four middle states as unresolved under the declared test, not as a winner ladder.
Product, engineering, and BI leaders Treat estimate plus margin of error as a bound pair through generation, semantic models, exports, and revisions; refuse requests that erase the uncertainty.
Visualization, HCI, and AI researchers Score source fidelity, uncertainty propagation, statistical decision, refusal usefulness, expert verdict, and reader outcome separately.
Educators and accessibility specialists Keep values, intervals, universe, unit, and comparison status available without color or hover; comprehension remains a human outcome, not a render check.

The interactive packet adds a second audience decision layer:

Audience Decision layer
Practitioners and editors After a filter changes, inspect the headline, marks, table, detail, shared URL, download, back button, and reset; a correct chart beside stale detail is still broken.
Product, engineering, and BI leaders Derive every surface from one canonical state, invalidate an ineligible selection atomically, and require replayable pilot receipts before treating a demo as application evidence.
Visualization, HCI, and AI researchers Start each trajectory from a declared isolated state and score visual fidelity, behavior, responsive delivery, keyboard equivalence, maintenance, and human usefulness separately.
Educators and accessibility specialists Require every filter, ticket, reset, and export without pointer input and announce state changes without stealing focus; actual assistive-technology use remains not run.

The manifest preflight changes how those audiences should read comparisons:

Audience Decision layer
Practitioners and editors Ask whether model, tools, context, token budget, retries, and reviewer access changed together before crediting one prompt or skill.
Product, engineering, and BI leaders Require an executable manifest, equivalent arm receipts, and total cost before pilot or procurement claims; a model label plus CLI version is insufficient.
Visualization, HCI, and AI researchers Report missing configuration before outcomes, keep a same-model bare diagnostic distinct from another local model, and preserve failures rather than imputing scores.
Educators and accessibility specialists Keep expert and human fields structurally present as not-run; missing keyboard, assistive-technology, learning, or comprehension evidence cannot be simulated.

The neutral runner turns the same answer into seven concrete decision products:

Audience Runner translation
Visualization and data practitioners Inspect the actual model-input manifest; an answer key, expected value, or acceptance rubric in context is part of the intervention.
BI and analytics leaders Require an exportable pilot directory with source, role, semantic context, artifact, renders, checks, failures, human state, and cost—not a screenshot plus score.
Visualization, HCI, and AI researchers Stage before inference, preregister the visibility boundary, hash outputs outside the arm, and report missing receipts as missing.
Data journalists, graphics editors, and newsroom developers Preserve the draft, source, renders, revision, and editorial judgment separately; do not let generation see the private answer key or let a summary erase failure.
Product, engineering, and tool teams Put neutrality in the filesystem: allowlisted input, isolated attempt roots, external output inventory, and a separate versioned provider adapter.
Educators, data-literacy, and accessibility specialists Deterministic checks can show that files or states exist; they cannot stand in for comprehension, transfer, keyboard experience, or assistive-technology use.
Executives, editors, and broad AI readers The evaluation plumbing is testable without spending. No model has run, and six model, budget, intervention, and acceptance decisions remain.

One layer means one controlled difference

A skill package is an installation and use surface, not automatically a causal unit. The two pinned public Vizier skills combine reader framing, form and pattern retrieval, honesty questions, encoding and palette guidance, deterministic commands, actual-artifact inspection, optional corpus retrieval, and critique composition. Testing either whole package would show the value of that bundle under its budget, not which mechanism mattered.

One candidate is now prepared but deliberately unselected: artifact-evidence reconciliation checkpoint v0.1. It enters once after the first executable artifact and neutral output inventory, before repair or finalization. For each decision-bearing visible claim or state, it records the artifact hash, available source/transform/query/calculation/branch/state evidence, and match, mismatch, or unresolved. Missing states remain missing; discrepancies are frozen before any bounded repair. Overall pass is valid only when every surface matches, no required state is missing, and no repair remains. The one failure target is silent divergence between the delivered artifact and the evidence behind it.

Use this five-part test before calling anything “one layer”:

  1. one named mechanism and one expected failure target;
  2. one lifecycle insertion point and one versioned output receipt;
  3. identical model, harness, inputs, tools, budgets, repair cap, and evaluator;
  4. evaluator answers and sentinel mappings kept outside model input; and
  5. an explicit owner decision distinct from a passing protocol verifier.

The same candidate changes seven audience decisions without changing its evidence status:

Audience One-layer decision
Visualization and data practitioners Ask which one checkpoint changed; do not credit a prompt when form, context, tools, retries, or critique changed too.
BI and analytics leaders Name the one control, insertion, wrong number or scope leak it targets, proof directory, false alarms, and review cost.
Visualization, HCI, and AI researchers Register the delta, exclusions, visibility boundary, common budget, and output schema; protocol readiness is not an effect estimate.
Data journalists, graphics editors, and newsroom developers Call a whole workflow a bundle; attribute one editorial mechanism only when it alone changed and the failed draft remains inspectable.
Product, engineering, and tool teams Version framing, retrieval, construction guidance, deterministic tools, artifact inspection, and critique composition as separable lifecycle hooks.
Educators, data-literacy, and accessibility specialists An evidence mismatch check can find inconsistent labels or states; it cannot supply learning, comprehension, keyboard, or assistive-technology acceptance.
Executives, editors, and broad AI readers One candidate is testable, not selected or proven. Owner ratification and the other five gates still precede comparison.

A pass needs the right authority

“Human review” is not one interchangeable box. A context reviewer decides whether an artifact supports the stated job. A domain expert checks semantics and interpretations that deterministic invariants cannot settle. A knowledgeable accessibility evaluator assesses a declared conformance scope. Relevant disabled users show how a delivered task works on their setups. Intended users accept or reject the artifact for the decision. A later independent maintainer supplies handoff and change evidence.

These roles cannot be filled by an automated critic. The W3C evaluation overview says no tool alone determines accessibility. Its guidance on involving users also says user evaluation complements rather than replaces standards work and must report participant scope without overgeneralizing.

The prepared E0 candidate makes that boundary executable:

Outcome lane Eligible authority Mechanical-only state
Contestable context review Named graphics/data editor, analyst, finance owner, stakeholder, or operations manager appropriate to the fixture not-run
Domain-expert review Named source, semantic-model, survey-methods, measurement, or operations expert appropriate to the claim not-run
Accessibility conformance Named knowledgeable accessibility evaluator not-run
Accessibility user evaluation Relevant disabled users, with task, setup, and participant scope retained not-run
Intended-user acceptance Named person in the fixture’s declared audience not-run
Maintenance handoff Later independent maintainer working from the retained artifact and history not-run

Five fixtures times six lanes yields 30 explicit not-run records. A scope receipt rejects a human pass without a named human authority and rejects an automated tool as that authority. A deterministic result may therefore be reported only as a mechanical pass for the named fixture and arm—not as unqualified acceptance, accessibility, usefulness, comprehension, maintenance, or Vizier efficacy. The research owner has not selected mechanical-only scope or named people, so this is a candidate policy with zero model calls and six remaining preflight blockers.

The audience consequence is concrete:

Audience Authority decision
Visualization and data practitioners Keep checks, editors, domain reviewers, intended readers, and later maintainers as separate stopping conditions.
BI and analytics leaders Name finance, semantic/governance, user, incident, and maintenance ownership before upgrading a synthetic mechanical result.
Visualization, HCI, and AI researchers Register authority, independence, information access, participant scope, and evidence; treat not-run as missing outcome data, not zero or pass.
Data journalists, graphics editors, and newsroom developers A palette collision can be computed; source binding, editorial acceptance, accessibility, and reader trust still need their own authorities.
Product, engineering, and tool teams Encode authority as typed receipt state and reject models or tools that attempt to fill a human lane.
Educators, data-literacy, and accessibility specialists Keep conformance evaluation and disabled-user experience distinct, and report the limits of each participant scope.
Executives, editors, and broad AI readers Five fixtures have one verified authority candidate; 30 human states remain unrun, no scope is selected, and no person was contacted.

Six answers, then one release

The six remaining E0 blockers are now one decision register, not one blanket approval. It offers 17 admissible paths across four authority roles and keeps each answer separate:

Gate Who can answer Advancing answer
Served-model identity Model provider or runner Supply an immutable snapshot and served receipt.
Server-harness custody Harness provider or runner Export the reasoning/output envelope, harness build, instruction digests, and tool-service versions.
Bare-model interface Provider for access; research owner for protocol Supply a same-model raw surface, or amend the protocol to drop that arm and narrow the claims.
Authorized budget Research owner Name the provider/model route, price basis, call/token caps, and stop condition, including for an evidenced zero-cost route.
One Vizier layer Vizier owner Ratify the exact v0.1 candidate; revision or rejection keeps the layer unselected.
Acceptance authority Research owner Ratify mechanical-only scope or require a real named-authority protocol before running.

An unavailable surface, declined budget, requested revision, or rejection is a valid recorded decision, but it cannot make the run ready. Even six advancing answers do not trigger a call: the research owner must separately release the hash of that resolved decision set against a named run plan, call cap, and stop condition. Decision receipts may identify a non-secret route and hash evidence; credential values never enter custody.

The verified candidate rejects wrong-role, incomplete, secret-bearing, partial, and six-answer-without-release states. That is protocol integrity, not approval. All six selections and the final release remain null, spend and readiness remain false, no person was contacted, and no model was called.

The audience translations make the same state useful without changing it:

Audience Decision-register translation
Visualization and data practitioners Six blank answers mean “not run,” not “almost passed”; ask which model, harness, budget, layer, and acceptance scope were actually selected.
BI and analytics leaders Use the register as a pilot approval sheet: provider identity, exportable envelope, cost cap, named control, acceptance policy, and accountable release.
Visualization, HCI, and AI researchers Publish gate receipts, amendments, and the decision-set digest; preserve unavailable surfaces and missing human outcomes rather than imputing them.
Data journalists, graphics editors, and newsroom developers Keep provider facts, editorial/product choice, reader acceptance, and publication judgment in separate hands.
Product, engineering, and tool teams Implement typed fail-closed state; provider adapters return non-secret receipts and never mutate owner choices or infer release from credentials.
Educators, data-literacy, and accessibility specialists Harness and budget approval do not establish learning, comprehension, conformance, or disabled-user experience; those lanes stay not-run until performed.
Executives, editors, and broad AI readers The experiment is prepared but not authorized: six decisions and one release remain, with zero calls and no efficacy result.

Score layers, not “quality”

The prepared E0 scorer has no overall score, winner, rank, universal pass, or efficacy field. It reports two mechanical tiers—computable checks and checkable obligations—plus six named nonmechanical lanes: context review, domain review, accessibility conformance, disabled-user evaluation, intended-user acceptance, and maintenance handoff. Cost stays beside those outcomes.

A mechanical tier passes only when it contains at least one check and all checks pass. One failure yields fail; an empty tier or any not-run yields incomplete. Each nonmechanical result retains its exact authority, rationale, scope, and evidence. not-run remains missing data, and one human role cannot sign another role’s lane.

The zero-call verifier created 15 synthetic fixture-arm ledgers and preserved 90 nonmechanical not-run records. It rejected unnamed and cross-lane human authority and injected aggregate score/pass fields. That establishes scoring plumbing, not 15 successful artifacts or a Vizier effect.

Audience Tier-separated decision
Visualization and data practitioners Read exact checks, contextual review, intended-user acceptance, and maintenance separately; a clean render is not a good outcome by itself.
BI and analytics leaders Require measure/permission checks, owner review, user acceptance, incident evidence, cost, and handoff as separate rows before adoption.
Visualization, HCI, and AI researchers Compare arms by predeclared lane, preserve missingness and parse failures, and report trade-offs rather than a composite leaderboard.
Data journalists, graphics editors, and newsroom developers Keep source/mark checks, editorial judgment, accessibility, reader evidence, and publication authority distinct.
Product, engineering, and tool teams Emit typed receipts from each evaluator and reject cross-lane authority or schemas that add an overall pass.
Educators, data-literacy, and accessibility specialists Conformance, disabled-user task evidence, learning, and transfer answer different questions and cannot substitute for one another.
Executives, editors, and broad AI readers “15 ledgers verified” means the scorer works on synthetic records; no AI output or human outcome succeeded.

Integrity and readability need a blinded test

R4 now has a preregistered pilot and five frozen inputs, not a result. Five opaque-ID SVGs, five neutral briefs, and nine source-evidence files make one 19-file model-input allowlist. Five context mappings and known-defect records stay in separately hashed evaluator-only custody. The model can inspect the same visual and evidence under either lens without receiving the answer key.

The two lenses now also have one hashed prompt and one fail-closed raw-response schema each. A response may complete with evidence-grounded findings, complete with no findings, or explicitly abstain. Twelve synthetic adverse records show that aggregate fields, mixed lenses, wrong identities, incoherent evidence, and findings attached to an abstention are rejected. This is tested local plumbing; it does not show that a provider supports the format or a model will follow it.

Three exact-model slots, the independent human data editor, calibrated tie rule, budget, release, and results are still empty. Future models produce separately blinded integrity and readability observations without scoring themselves; the editor later scores them against hidden defect truth and a declared readability rubric. Only the two within-lens partial orderings are compared. They never become one quality score.

The remaining approval path is now explicit. Five gate categories become seven attributable receipts: one exact served receipt for each of M1, M2, and M3; editor-signed role acceptance; same-editor calibration; a research-owner capped budget; and a final research-owner release over the exact six- prerequisite digest. The budget must cover 30 primary records—three models by five stimuli by two lenses—with retries capped separately.

Twenty advancing and non-advancing receipt paths validate. Sixteen synthetic adverse states show that wrong authority, cross-subject choices, incomplete or unsafe evidence, duplicate/partial prerequisites, a refusal, missing release, wrong digests or hashes, different editors, route mismatch, call-cap mismatch, and duplicate model identity remain blocked. This tests the handoff mechanism; all six real prerequisites are pending and release is false.

The editor’s future calibration task is now prepared but blank. It contains ten baselines—five stimuli under each lens—and ten synthetic response-quality vignettes, including two A/B-swapped repeats. The response can complete, request revision, or state that usable resolution cannot be established; fourteen adverse custody, coverage, repeat, state, and aggregate-result records fail. No editor, appointment, calibration, resolution, or tie release exists.

Custody layer What is present What it licenses
Model input 5 visuals + 5 briefs + 9 evidence files, each path and hash allowlisted A future evidence-grounded audit after provider and owner release
Evaluator only 5 defect mappings, expected findings, false-allegation guards, and 9 exact origin bindings Later scoring by the named editor; never model context
Prompt/output 2 lens prompts + 2 schemas; finding, empty, and abstaining states Inspectable raw observations, not self-scores or a model comparison
Severe checks Duplicate IDs, truth-path leakage, wrong hashes, and missing evidence are rejected Confidence in staging boundaries, not model quality
Response checks 12 aggregate, mixed-lens, identity, evidence, and state failures are rejected Confidence in the response boundary, not provider compliance
Visual inspection Corrected label overlap, an undeclared truncated baseline, and hidden export text before freeze A cleaner pilot input set, not an audience or accessibility result
Receipt Who supplies it Current state
M1, M2, M3 served routes Provider or runner for each exact slot All three pending
Editor acceptance The named independent data editor Pending; no person contacted
Editor calibration The same accepted editor, with retained evidence Pending; resolution and tie behavior remain null
Capped 30-primary-call budget Research owner Pending; no spend authorized
Exact-set release Research owner, naming the six-receipt digest and frozen hashes False; no call authorized
Audience R4 decision
Practitioners Ask for a bounded observation, exact visual/source locator, and scope; an empty finding set or abstention is evidence, not a formatting failure.
BI leaders Require completed integrity records to consult both the visual and source evidence; one fictional semantic model is still not a procurement score.
Researchers Preserve raw lens records, empty findings, abstentions, and parse failures; apply one human ordering later instead of collecting model self-scores.
Newsrooms Models produce cited observations only. A named data editor retains hidden-truth, false-allegation, score, tie, and publication authority.
Product/tool teams Validate returned JSON against the frozen lens schema and exact stimulus allowlist; record provider support in a served receipt rather than assuming it.
Educators/accessibility specialists Readability observations are not comprehension, learning, conformance, or disabled-user results; integrity observations are not reader performance.
Broad AI readers Two prompts and two schemas pass synthetic tests; zero models or people have answered them, so no winner or divergence exists.

The same register changes the decision question by audience. Practitioners can ask whether all seven receipts are inspectable. Analytics leaders can use them as a pilot-control sheet. Researchers can retain refusals and provider failures as results about the protocol rather than silently changing it. Newsrooms keep provider, adjudication, research release, and publication authority separate. Product teams can re-hash a route adaptation instead of hiding it in the runner. Education and accessibility readers can see that editor calibration is not learner or disabled-user evidence. Broad readers get the shortest honest status: the approval path is testable, but every real answer is blank.

For practitioners and analytics leaders, the additional check is whether the editor wrote the baseline before seeing model output. Researchers should treat the swapped repeats as a consistency check, not inter-rater reliability. Newsrooms retain editorial/publication authority; product teams keep the packet evaluator-only; education and accessibility readers should not read review- scale calibration as comprehension or disabled-user evidence. Broad readers can say only that the editor now has a concrete future task.

What can and cannot be projected

The observed direction supports three bounded forecasts:

  1. Likely to commoditize: syntactically valid chart code, conventional chart selection, basic styling, routine repair, and generic “inspect your render” advice. These should be short, replaceable defaults, not a large permanent application-specific doctrine.
  2. Likely to migrate into models or harnesses: generic planning, self-review, browser/tool use, and library navigation. A visualization system should consume these when they work and avoid duplicating their orchestration.
  3. Unlikely to be solved by baseline capability alone: the local question, audience, semantic definitions, authoritative data, denominator, source custody, consequence, publication boundary, environment-specific evidence, and real-reader outcomes. These are facts and authorities the model does not acquire merely by becoming better at code or vision.

Linear projection would be false precision. The practical forecast is that the value of generic instruction will decay fastest, the value of executable local evidence will persist, and the value of context and authority boundaries will increase as agents become capable of taking more consequential action.

How evidence was graded

The same word—“works”—covers incompatible outcomes in this literature. This report keeps seven layers separate:

Layer What can be established Typical evidence What it does not establish
Structure Required fields or files exist schema validation, AST checks, spec comparison execution, correct values, usefulness
Execution Code compiles/runs and a nonblank artifact appears sandbox, renderer, browser, console/network checks correct binding or interpretation
Data fidelity Values, aggregations, filters, and joins match source data deterministic recomputation, exact comparisons, query receipts legibility or audience value
Visual integrity Encodings are coherent and nondeceptive mark/encoding checks, render inspection, targeted critic ease of reading or insight
Interaction Controls and linked views behave as intended browser replay, action coverage, before/after state that the interaction supports a good analysis
Analytical support The artifact helps address a task task-grounded comparison, expert study, realistic workflow learning, retention, or downstream decision quality
Reader outcome A defined audience understands, remembers, or decides better controlled human study with audience/task measures generalization beyond that population and context

Passing a lower layer is a prerequisite, not a proxy for all higher layers. “Rendered,” “valid,” “high similarity,” and “preferred by a vision model” are different claims.

Evidence labels used below:

Technique map

Technique Representative systems Best current support Environment fit Failure boundary Current evidence-based posture
Prompt-only visualization guidance claude-skillz data-visualization, Markdown Viewer Vega skill C/D: inspectable instructions; no behavioral comparison found one-off, low-risk chart creation where the renderer is already known prose can be stale, ignored, or internally inconsistent; no data or render proof Reject as a sufficient workflow; retain only as baseline
Versioned library constraints plus retrieval AntV G2/G6/X6 skills C: large reference corpus, retrieval datasets, structural checks, included generation results coding against a fast-changing visualization API reported “success” can mean response returned; saved results were not browser-rendered Try as a retrieval intervention, not as quality proof
Chart contract before implementation OpenAI visualize-data; Vizro design specs C: explicit contracts, runtime tests, required specs and receipts coding agents, reports, dashboards, multi-surface delivery instruction compliance is probabilistic; contract can document a wrong question Borrow now
Structured analytic or visual IR NL4DV-LLM, nvAgent VQL, NL2Dashboard IR, RaivenDSL, ViviDoc SRTC, Flint A/B in bounded tasks plus ViviDoc’s same-team C-grade ablation repeated generation/editing, scientific views, dashboards, interactive documents, renderer portability schema can exclude useful forms; semantic binding can still be wrong; creator satisfaction is not reader outcome Borrow the principle; test the smallest local IR
Deterministic compiler/renderer Raiven, NL2Dashboard, Flint, Vega-Lite-based tools Raiven B/A in fully specified reproduction; NL2Dashboard B scientific visualization, stable dashboard grammar, controlled environments does not discover the analytical question; expressiveness ceiling and compiler defects remain Try against direct code at equal task/budget
Direct manipulation plus natural language Data Formulator 2, DashChat, Tableau Agent Data Formulator 2 and DashChat A within small studies exploratory analysis and prototype negotiation user studies are small; current products have evolved beyond published evaluations Borrow interaction pattern
Persistent history, branching, reversion Data Formulator/Data Threads, DashChat history A/C: observed in studies and current implementation nonlinear analysis and stakeholder iteration provenance can record bad branches without detecting them Borrow now
Semantic-layer grounding Looker/LookML, Power BI semantic models, Tableau field metadata; current Data Formulator connectors C/D: strong mechanism and governance rationale; limited public comparative outcome evidence governed enterprise BI with established models semantic layer can be stale or wrong; unmodeled questions remain hard Borrow as an input contract when one exists; never infer semantic truth from availability
Execution-guided repair nvAgent, PlotGen lexical loop, Lumen agents, Vizro testing nvAgent B; other evidence mixed syntax/API-heavy generation and multi-table queries fixes what throws, not silent semantic error; repeated loops add cost Borrow bounded repair after typed failure
Rendered-image critique VisJudge, Raiven VMPC judge, PlotGen visual agent, NL2Dashboard critic VisJudge B on its expert-adjudicated quality rubric; Raiven B with human correlation; other evidence weaker visible composition, readability, mark, label, and bounded perceptual failures a screenshot critic cannot verify source fidelity, interaction, responsive states, or reader outcomes Try only as one evidence channel
Chart-specific perception and tools ChartAgent, ChartREG++, chart parsers, OCR/layout models Bounded specialist gains in numeric QA, chart-mark grounding, and document parsing; current general models lead some new transfer tests exact extraction or localization failures that a general critic cannot resolve reliably parsing is not critique; tools still fail; specialist benchmark fit can decay quickly Route to a measured component failure, not to “quality” in general
Deterministic data/integrity checks Vizro aggregation/color scripts; Raiven data-hallucination checks; local regression guards B/C and strong causal fit any generated chart with inspectable source data check coverage is necessarily partial Borrow now and expand by failure class
Browser and interaction evidence Vizro Playwright workflow, DashArena executor DashArena A/B; Vizro C interactive dashboards and browser-delivered reports scripted path may miss latent behavior; clean execution can hide data errors Borrow now for interactive surfaces
Model-authored replayable analytical trajectory DashArena A/B: 234 tasks, human-calibrated pairwise judging open-ended interactive dashboard generation trajectory is partial and model-authored; benchmark is Tableau-seed-biased Try as an evaluation receipt, not as product behavior
Multiple role agents LIDA, PlotGen, nvAgent, Lumen, DashChat, NL2Dashboard mixed: nvAgent/DashChat positive, equal-budget general evidence negative, several ablations nuanced heterogeneous access/tool boundaries or parallel independent work role theater, correlated critics, token/latency growth, weak baselines Use only where the role owns a distinct capability or evidence boundary
Synthetic reader/persona panels LIDA goal personas; broader synthetic-user literature weak for audience validity; negative evidence on faithful simulation brainstorming possible questions only false confidence about real readers, difficulty, aesthetics, and persuasion Reject for acceptance or audience claims

Research and implementation landscape

This landscape shows where research techniques appear in working systems, agent-skill packages, and governed BI products. The entries are examples and evidence records, not product scores: a paper can support a bounded performance claim, a repository can expose an implementation, and product documentation can describe a contract, but those are different kinds of evidence.

Full tools and research systems

Offering Actual mechanism Evidence read Assessment
Data Formulator Natural language handles transformations and agentic exploration while a GUI handles explicit visual encodings; visible tables, code/explanations, threads, branches, and current Flint semantic specs support refinement. Data Formulator 2: eight participants reproduced 16 charts and 12 nontrivial transformations; all completed the tasks, with distinct branch/depth strategies. Current 0.8 alpha is substantially broader than the studied 2024 prototype. A/C. Strongest authoring interaction pattern. The user study supports learnability and verification behavior, not long-term analytical correctness or the efficacy of the 2026 agent stack.
Raiven The model sees metadata and emits RaivenDSL; a deterministic compiler creates coordinated 2D, 3D, and tabular views with controls. 100 fully specified prompts; 100% compile and .988 VMPC versus .800–.867 compile and .678–.721 VMPC for direct-code baselines. Seven visualization experts completed three replication tasks; five preferred Raiven and two had no preference. A/B, bounded. Compelling for unfamiliar scientific/3D prototyping. Most gains are in SciVis; InfoVis baselines nearly match. Tasks reproduce specified views rather than discover questions. Authors were the graders, and the tool can still encode logically misleading order/null choices.
NL2Dashboard A compact IR separates analysis/content/layout from deterministic rendering; atomic modification operators avoid rewriting a dashboard; planner, coder, and optional visual critic operate around the IR. Ten tables across finance, education, and government; seven modification classes. It scored 11.89/15 in generation and 11.93/15 in modification under an LLM judge, completed all edit tasks, and used much lower output-token/dashboard ratios than web-product baselines. B, promising but not decisive. The IR/edit result is strong. The comparison uses different model interfaces/backbones, only ten source tables, an LLM quality judge, no user study, and no equal-budget control. The paper reports diminishing critic returns and recommends zero or one round.
ViviDoc A DocSpec decomposes interactive units into State, Render, Transition, and Constraint before code; authors can edit the plan, style, and result. 101 topics across 11 domains; a same-pipeline ablation across three backbones; 36 blind-rated outputs; and 12 people authoring two documents each. The largest reported interaction-quality lift from DocSpec was 41%. C, useful authoring mechanism. Same-team evidence and partly model-judged metrics. DOM change and creator satisfaction do not establish source truth, reader learning, accessible or production delivery, maintenance, or autonomy.
nvAgent Processor filters/augments database schema, composer uses sketch-and-fill to make VQL, validator translates to Python and iterates on execution errors. ACL 2025; VisEval has 2,524 NL/visualization pairs over 146 databases. With GPT-4o it reached 85.63% single-table and 81.07% multi-table pass rate, +7.88 and +9.23 points over the best baselines. B. Supports structured planning and execution validation, especially for multi-table work. The composer carries most of the ablation gain. Removing the processor slightly improved the GPT-4o average while hurting weaker/multi-table settings. Five paper authors performed human annotation; the authors acknowledge evaluator bias, temporal errors, and incomplete semantic metrics.
DashChat Industrial-dashboard pattern retrieval, a DSL, intent-specific parallel agents, explicit evaluation/repair, chat plus structured edit bubbles, and visual history/reversion. Fifty held-out prompts: 100% executable, 94% exact spec consistency, and 41.4 s end-to-end; the two single-pass baselines reached 76–80% consistency and took 62.9–134.7 s. User study: 17 domain professionals and 11 designers; designers compared it with Tableau after a 30-minute tutorial. A/B for rapid prototypes. Strong evidence for a constrained prototyping environment and iterative negotiation, not production analytics. It generates mock data; participants identified domain mismatch, acronym, direct-editing, and style-control limits. The Tableau comparison favors a prompt-first prototype task and says little about governed, production data work.
Lumen Coordinator routes to SQL, Vega-Lite, Deck.gl, chat, source, table/document-list, and validation agents over serializable declarative pipelines and views. Deterministic profiling and cleaning precede some agent work. Current, inspectable implementation and tests; no persuasive published comparative efficacy study found. A notable validation path fails open by assuming completion when structured validation cannot be parsed. C. Strong mechanism source, not outcome evidence. Borrow serializable pipelines, real access-based roles, and profiling-before-aggregation. Reject fail-open semantic completion.
LIDA Data summarization, persona-conditioned goal generation, chart code generation, six-dimension LLM evaluation, repair, and recommendation. Influential open-source baseline; inspected revision has not moved since 2024-08. The evaluator is code/text-prompt based and repair echoes feedback into another generation call; it does not inspect the final rendered image or produce regression evidence. C/D and historically useful. Keep as the canonical prompt-pipeline baseline. Reject its evaluator/persona pattern as current acceptance evidence.
PlotGen Query planner and code generator followed by numeric, lexical, and visual feedback agents; the numeric agent de-renders the result with a VLM. MatPlotBench 100: 65.67 with GPT-4 versus 61.16 for MatPlotAgent and 48.86 direct. Five participants reviewed 200 sampled requests; only 40.5% were completely accurate and 24.5% somewhat accurate. B-, directionally useful. Supports multimodal feedback, especially lexical/visual checks, but lacks an equal-budget baseline, relies heavily on GPT-4V throughout, and contains reporting inconsistencies. It does not establish that multiple agents are the cause.
NL4DV-LLM The model emits an inspectable analytical specification (attributeMap, taskMap, visList) and can return multiple interpretations for ambiguous prompts. 740 queries over three datasets: GPT-4 prompt approach 87.02% versus rule-based NL4DV 64.05%, at roughly 25 s versus 3 s. Two authors graded, with a third tie-breaker; any valid ambiguous interpretation could count. B, older-model evidence. Borrow explicit task vocabulary, mapping visibility, and multiple interpretations. Do not trust its generated confidence scores or equate valid syntax with correct attribute/encoding binding.
Data Formulator 2 Concept binding: direct manipulation specifies encodings; concise natural language requests missing transformations. Data threads preserve branch/backtrack context. CHI 2025 study described above. Participants used charts, transformed tables, code, and explanations differently to verify outputs. A within a small reproduction study. This is better evidence for mixed initiative than for agent autonomy.

Skill and plugin packages

Package Package class Verification actually present Outcome evidence Assessment
Vizro end-to-end flow Six skills split design, chart/layout selection, build, YAML, and actions; five required spec/test artifacts AST checks for raw/unaggregated charts and color policy; required terminal inspection; Playwright walk of every page and every action; console, network 500, server traceback, screenshot/spec comparison, and test report Three dashboard-build and four interaction eval prompts with explicit expectations; README says tested with two Claude 4.6 models, but no aggregate held-out result or independent user outcome found C, strongest inspected skill mechanics. Borrow staged specs, targeted deterministic checks, action enumeration, browser evidence, and test receipts. Do not infer general chart quality from seven fixtures.
OpenAI visualize-data Large workflow skill: question/takeaway first, chart contract, data sufficiency thresholds, surface routing, denominator/uncertainty/source rules, final-context rendering and inspection Repository tests cover renderer, transform, tooltip, axis-domain, HTML fallback, and delivery contracts; the skill requires QA in the delivered surface No held-out behavioral comparison of an agent with/without this skill found C. Strong contract and coverage checklist. Borrow the chart contract and final-context QA. Treat prose thresholds as revisable defaults, not universal laws.
AntV chart visualization skills Thin chart-image API skill plus deep G2/G6/X6 skills with strict version constraints and hybrid retrieval over a large reference corpus Eval code contains structural/API checks, code similarity, a Playwright render tester, blank detection, and a VLM visual scorer Included saved runs cover 174 G2, 97 G6, and 136 X6 cases with high structural similarity/hit rates. The July retrieval result files do not contain render results; “success” largely means generation completed without recorded structural failure. No no-skill baseline is included in those files. C. Excellent evidence that version-specific constraints and progressive retrieval target real library hallucinations; insufficient evidence of visual correctness. Run a local with/without retrieval-and-render ablation before transfer.
SciVisAgentSkills Version-pinned operational guides for napari, ParaView, Topology ToolKit, and VMD/MDAnalysis, including headless execution and render–inspect–adjust loops Deterministic image, code, and rule checks plus multimodal judging across 108 expert-designed tasks Claude Code and Codex were each tested three times with and without the relevant skills. Quality improved in all ten suite-by-agent comparisons, although one completion measure fell. B, direct but bounded. This is the clearest visualization-specific skill ablation found. The authors built and evaluated their own packages; the tasks are scientific, and no independent reproduction or reader outcome was found.
Markdown Viewer Vega skill Compact renderer adapter with syntax notes and examples No task fixtures, data checks, or render loop found in the skill None found C/D. Useful as surface syntax, not a visualization method.
data-visualization in claude-skillz Long single-file primer covering chart selection, Cleveland–McGill ordering, accessibility, performance, libraries, and layout No scripts, fixtures, or evaluation harness found None found D. Good checklist specimen and prompt-only baseline. It packages advice but cannot show that an agent followed it or that a reader benefited.

These packages make a useful maturity ladder:

static advice
  -> versioned constraints and on-demand retrieval
  -> explicit design/build artifacts
  -> deterministic source/code checks
  -> final renderer/browser inspection
  -> interaction coverage and evidence receipts
  -> held-out behavioral comparison with and without the package
  -> human outcome study

SciVisAgentSkills reaches the paired-ablation rung for specialized scientific work, but not independent reproduction or a human outcome. Vizro reaches furthest on execution evidence; AntV has the largest included retrieval/code benchmark; the OpenAI skill has the broadest chart-contract and delivery QA. Those are different strengths and should not be collapsed into an install count or one “best skill.” The companion skill-package deep dive compares 18 files or families and separates registry installations from repository popularity and behavioral evidence.

Governed commercial environments

Product surface Current first-party mechanism What it suggests Evidence limit
Power BI Copilot Builds a report page by selecting tables, fields, measures, and charts from a semantic model; generated pages remain editable with normal tools; answers can reference source visuals. Bind generation to governed measures and retain direct author control. Current capability documentation, not a comparative accuracy or user-outcome study.
Tableau Agent Works within a connected data source and current worksheet state; uses field metadata and sampled values; creates/changes visualizations, calculations, filters, and sorts; dashboard Q&A entered beta in July 2026. Keep agent scope close to existing authoring state and make direct manipulation the recovery path. Tableau explicitly says to review results and treats the output as a starting point. It currently cannot choose a source, model data, build full dashboards in viz authoring, or create many interactions.
Looker Conversational Analytics Grounds queries in LookML, permissions, descriptions, samples/fuzzy value search, custom instructions, business glossaries, and optional verified queries. The agent selects fields/filters while Looker composes database queries; optional Python handles advanced analysis. A maintained semantic layer and verified examples are more reliable grounding than a generic analyst persona. Different domain agents can be policy/configuration packages over shared governed data. Product documentation says outputs can be plausibly wrong and must be validated. No public evidence read here isolates which grounding feature improves end-user decisions.

The commercial systems are especially environment-dependent. Their main advantage is not a universally better model. It is custody of semantic models, permissions, field metadata, verified queries, authoring state, and the native renderer. That advantage does not transfer to an open-file editorial workflow unless another system is given equivalent data contracts.

The evaluation frontier

DashArena materially changes the map

DashArena, published 2026-08-11, contains 234 open-ended tasks derived from high-quality Tableau Public dashboards across 14 clusters. A candidate returns a single-file ECharts dashboard and a structured two-turn interaction trajectory. A Playwright executor replays the trajectory and gives the pairwise judge task context, screenshots, schema, and execution evidence.

The strongest results are about the evaluation method:

This qualifies the earlier conclusion that communicative-value evaluation was empty. DashArena now provides a serious, human-calibrated measure of task-grounded analytical support and interaction quality. It does not reverse the conclusion about readers. The benchmark does not observe whether a target audience comprehends an argument, learns a concept, remembers the message, or makes a better real decision. Its pairwise aggregate is also not a universal taste or audience model. The right update is “a major middle layer now exists,” not “communicative value is solved.”

The naming has outrun the instrument

Two further 2026 benchmarks reach toward communicative value and then measure something else. Reading them together is more informative than reading either alone, because they fail in the same direction from opposite starting points.

SciVisAgentBench is the most systematic scientific-visualization agent benchmark released so far: 108 expert-crafted cases over a four-dimensional taxonomy, a multimodal evaluation pipeline, and an IRB-approved validity study with 12 domain experts. Its taxonomy makes Scientific Insight Derivation a first-class visualization operation. But that is a task the agent performs, not a property of the artifact that is scored. Scoring is rubric agreement with a single expert-authored ground-truth visualization, judged 0–10 by an MLLM, plus PSNR, SSIM, and LPIPS similarity to that reference, plus token and time efficiency. Tasks are deliberately curated to admit one explicit outcome. The 12 experts are not readers being measured; they are raters whose agreement with the LLM judge is being measured. The paper is candid that open-ended goals with multiple valid outcomes “remain difficult to evaluate reproducibly at a benchmark scale,” and names the same obstacle from the inside: different visualizations may convey the same insight, which complicates automated scoring.

MultiVis-Agent goes further in language and no further in instrument. Its evaluation has an explicit high-level perceptual layer of six weighted dimensions — chart-type appropriateness, spatial layout, textual elements, data representation, visual styling, global clarity — described as “capturing human-perceived effectiveness.” The scorer is a vision-language model applying rubrics to a rendered image. No human reader appears anywhere in the evaluation.

The gap, stated precisely, is no longer that the field has failed to name communicative value. Both of these name it, or something adjacent to it. The gap is that naming has outrun measurement, and a model judge scoring a rubric is increasingly being asked to stand in for a reader. That substitution is convenient, reproducible, and cheap, and it is not the same evidence. A benchmark can report a high perceptual score for a chart that no human has read.

For anyone selecting a benchmark: read what the metric is computed on, not what the metric is called. A dimension named for clarity, effectiveness, or insight may be a model’s rubric score against a reference image. That is a useful signal about conformance and a weak one about communication.

Dashboard2Code exposes state that a screenshot can hide

Dashboard2Code reconstructs interactive Plotly Dash applications from screenshots, optional DOM, and interaction. Its benchmark combines 58 real-world seed dashboards with 122 generated examples, then checks visual fidelity, code, dynamic browser behavior, and 450 interaction tasks. Ninety generated dashboards were also scored by three visualization- experienced graduate evaluators; the final automatic metric correlated .781 with their ratings.

The most important result is not one total score. DOM access sharply improves exploration, complex callbacks remain harder, and systems can produce a visually responsive but factually wrong state. A static image judge cannot see that failure. The scope boundary matters just as much: Plotly Dash only, fixed desktop viewport, no animation or popups, and no reader study.

Chartography makes professional reading a separate capability

Chartography, released 2026-08-11, contains 100 difficult professional chart-reading tasks across 12 domain labels. Each task was authored by a practitioner and independently verified by three experts, including an acceptable answer range. The best of 30 frontier configurations reached 45.0% mean pass@1 over twenty trials per task.

Higher reasoning effort usually helped, but not enough to erase the main failure: sparse axes, 3D projections, contours, and domain conventions can be misread before reasoning begins. More tokens then produce a longer explanation of the wrong visual premise. The benchmark is adversarially difficulty- screened, so it does not say models fail on 55% of ordinary charts. It does say that a general “chart literacy” score is too broad a release gate for professional, consequential reading.

Raiven narrows generation failure in a particular environment

Raiven’s .988 VMPC is strong counterevidence to treating direct-code chart generation failure rates as immutable. A formal DSL plus compiler can nearly eliminate many specified-mark, encoding, linking, hallucination, and execution failures. But the benchmark fully specifies the target visualization, and the largest delta occurs in 3D/scientific tasks where generic web code generation is weak. In ordinary 2D information visualization, direct frontier models nearly match Raiven on its own metric.

The transferable claim is conditional: when an environment has a stable visual grammar, recurring hard syntax, and deterministic rendering, invest in the representation/compiler. It is not evidence that a DSL can choose the right question, identify a misleading comparison, or replace final human judgment.

Benchmark construction is part of the technique

The companion benchmark crosswalk documents the input object, requested output, source distribution, grader, human context, supported claim, and nonclaim for the benchmarks carrying the main conclusions in this review. It reveals a measurement progression:

  1. exact or near-exact code/spec matching;
  2. syntactic legality and render success;
  3. data/encoding correctness against a reference;
  4. chart extraction, visual grounding, question answering, and comparison;
  5. integrity or quality judgment with human calibration;
  6. interaction state and task-grounded replay;
  7. actual user performance, comprehension, or decision outcome.

Many impressive percentages in tool repositories live at levels 1–2. nvAgent and Raiven reach level 3 in bounded ways. ChartQAPro, Chartography, ChartDiff, and the two 2026 PolyChartQA benchmarks expand level 4 across realistic, professional, comparative, multi-chart, and multilingual cases. Misviz, VisJudge, and misleading-chart robustness studies measure parts of level 5. Dashboard2Code and DashArena reach level 6. Data Formulator 2 and DashChat provide small, contextual evidence at level 7 for authoring/prototyping experience, not for reader comprehension or decision quality. No one result spans the stack.

Three construction effects now deserve the same attention as model scores. Human-authored multi-chart questions were up to 27.4 percentage points harder than model-generated questions; ChartDiff’s lexical-overlap metrics favored systems that its human-aligned judge rated much worse; and misleading-chart system rankings changed between controlled synthetic and heterogeneous real charts. A benchmark is a task, distribution, and grader—not a neutral name.

Five validity threats block a configuration verdict

The latest audit asks a narrower question than the landscape: does this evidence authorize one Vizier configuration? The answer is no. It records five distinct boundaries rather than hiding them in one confidence label:

Threat Current result What must happen next
Cross-literature join retained unresolved test the visualization-specific architecture under equivalent current receipts
Flattering-prior risk partly tested, retained preserve the hostile pass’s failed hypotheses and preregister rivals before inference
Benchmark-to-quality inference rejected require a new receipt at every later evidence layer
Recency boundary bounded keep durable task constructs separate from perishable model/harness results
Single-configuration limitation retained unresolved run the comparison; prepared packets and verifiers are not results

Two direct human/model comparisons make the benchmark boundary concrete. CHART-6 evaluates eight models on 851 items from six visualization-literacy assessments, ten times each. The tested models underperformed the held human data on average, and every item-level model error pattern remained below the human noise ceiling. A newer high-level interpretation study compares 24 people’s descriptions with three multimodal models over 60 charts. Models matched the study’s bounded analytical designer intent more often, yet tended to enumerate structure and values where people more often built narratives and used visual scaffolding.

Those findings are not inconsistent once task, model, prompt, scoring, and vintage remain attached. They show why aggregate task success, human-like errors, analytical-intent matching, and reader experience are separate claims. The public rule is therefore strict: a benchmark licenses only its named task-and-grader lane. Overall quality, audience fit, comprehension, accessibility, accepted delivery, maintenance, and a configuration winner each require new evidence.

Research gap ledger

An open question is not the same as an empty field. Some gaps now contain a bounded study, some contain evidence that a failure occurs, and some still lack the measurement needed to make a decision. The status language below is literal:

Gap What the evidence establishes now What remains missing Next evidence that would change the status Status
One-pass competence on realistic work RealChart2Code and DashArena show that frontier systems can produce materially faithful charts and open-ended interactive dashboards, while still missing render, replay, and semantic requirements. The share of ordinary editorial, operational, scientific, and governed-BI work that is acceptable without repair. A stratified held-out task set sampled from real work, scored through data, render, interaction, and human acceptance. Measured but bounded [Q9]
Capability-to-practice bridge Ten primary cases were audited for technical-to-human joins. HAIChart links recommendation performance to 17-person controlled use, interactive task decomposition links system error to correction across 108 analyst episodes, and ChartAttack carries generated attacks into a 48-person reader experiment. An immutable system/version plus frozen acceptance contract and real delivery. No audited case continues through a complete reader or decision outcome, later maintenance, and whole cost. Carry one controlled result forward without changing artifacts, tasks, configuration, or version: accept under a frozen rule, deliver it, observe representative use, and return after a material event with full cost. 3 controlled partial · 0 accepted-delivery
Data and semantic fidelity nvAgent, Raiven, Text2Vis, and DashArena test bindings, calculations, or task-grounded semantics in bounded environments. DashArena found semantic defects even among execution-clean outputs. Accuracy against organization-owned measures, ambiguous fields, changing sources, permissions, and unstated local rules. Public field evaluations stratified by semantic-model quality, task ambiguity, and data conditions. Measured but bounded
Correction without regression RealChart2Code demonstrates regressive editing: a requested fix can introduce new faults in previously correct code. The novice study found 11 of 16 clutter fixes and 8 of 9 unusable-chart repairs failed. Accepted-artifact correction cost, abandonment, regression after delivery, and reliable stopping rules across models and environments. Multi-turn studies that retain every state, classify introduced and repaired faults, and follow work through acceptance. Failure demonstrated [Q15]
Critic architecture and specialized vision VisJudge-7B outperformed the tested general models on its expert-adjudicated quality rubric; ChartAgent’s chart-specific tools materially improved the same general reasoner on numeric QA; chart-aware grounding and OCR models add useful perception. New transfer tests also show strong general models overtaking older chart specialists. A routing rule, same-generator and equal-budget end-to-end critique lift, cost, interactive and mobile coverage, source-fidelity checks, independent reproduction, and reader outcomes. Evaluate deterministic checks, current general vision, narrow specialists, and composed critics on the same generated artifacts and consequential defect set. Partly answered [Q7] [Q16]
Readability versus integrity Frontier-model visualization literacy and misleading-chart detection can rank differently, showing that decoding and integrity are distinct capabilities. Whether that divergence reproduces on current generated artifacts and predicts actual human misreadings. A shared chart set scored independently for data truth, deceptive encoding, readability, and reader outcomes. Open [Q8]
Learning, authoring outcomes, and expertise The human-skills synthesis anchors expertise in consumption, construction, critique, and connection plus data, domain, tool, situated-judgment, and delivery resources. One randomized study supports immediate post-removal comprehension from proactive scaffolding; adjacent trials show assisted performance can outrun learning; a no-AI visualization study measures ordinary decay over 12 months. A protocol now crosses conventional, answer-oriented, and metacognitive practice with immediate withdrawal, six-week and six-month unfamiliar transfer, and a separate blinded reader stage. A stable definition of “novice”; approved and powered crossed capability/accessibility cells; a frozen model and harness; delayed visualization construction and far transfer; real-reader delivery on own devices and assistive paths; and effects over weeks or months. No captured longitudinal study causally estimates AI-driven visualization learning or atrophy. Approve, register, and run the protocol while keeping access, assisted performance, learning, cost, accessibility, and reader outcomes separate. A protocol verifier is not a human result. Partly answered; study designed, not run [Q13]
Reader comprehension and decisions A ten-row AI-role audit finds one direct controlled effect where selected AI-generated ChartAttack outputs reach independent readers. Lexara adds a distinct longitudinal middle: six CVA developers used a deployed evaluation toolkit for two weeks across 38 experiments, 57 newly authored cases, ten models, and six prompts. Three production afterlives and the later same-feature InfoViP field-use and operating line bring the audit to 15 named cases. Zero case binds an immutable artifact/build and exact configuration to accepted intended-reader delivery, consequential decision or calibrated trust, whole cost, and later post-release same-lineage recheck. Development selection, a bounded voluntary field evaluation, and at-scale processing do not fill those states. Keep twelve receipts on one lineage from AI role through artifact/build/configuration, acceptance, delivery, context, decision, trust, elapsed event, recheck, maintenance authority, and cost. Carry one evaluation or operational record into a frozen release and affected-audience return rather than borrowing receipts across cases. 15 named · 1 longitudinal evaluation · 1 counted field-use + operating near-miss · 0 full episodes [Q10] [Q14]
Decision-visualization review lineage A 2025 PRISMA review searched five databases through 1 July 2024 and retained 127 studies: 118 empirical papers and nine reviews. Its supplement prints 126 citations plus ?; the source-internal B59 overlay repairs that placeholder. A complete six-domain crosswalk now binds 122 of the 127 supplement positions to distinct A208 review keys. Of those keys, 114 carry DOIs, two carry confirmed PubMed identifiers, five expose no publisher identifier, and one carries a printed PMID that resolves to another paper. Five printed rows—Jiao (2022), Burt et al. (2017), Theis et al. (2018), M. Lu (2020), and Islam et al. (2022)—have exact-looking external DOI candidates but no review-controlled join, so none enters the authoritative frame. The conflicted B13 PMID also remains quarantined. Every lifecycle field is still unaudited; the qualifying count is unknown, not zero of 127. Preserve source label, authoritative review key, external candidate, identifier validation, and admission authority separately. Obtain review-author, publisher, extraction, or A208-controlled joins for the five gaps, repair or remove the conflicted PMID, and only then audit the stable frame against the eight receipts—or capture one direct same-artifact GenAI episode independently. 127 positions · 122 review keys · 5 authority gaps · denominator unknown
Artifact custody and audience delivery A bounded title-cue screen over the 122 authoritative A208 keys surfaces one explicit GPT study. Beyond Generating Code evaluates 91 quiz questions and nine homework assignments; its one-commit supplement pins 1,271 blobs spanning prompts, outputs, generated code, and grader bundles. The runner exposes model aliases rather than an immutable provider snapshot and complete configuration. The paper does not directly evaluate the final project. Government, UN, and primary-school audiences appear in prompts, not as recipients or evaluators. No named release, later commit, audience delivery, outcome, or recheck is exposed. Carry the pinned bundle into a frozen release and acceptance rule, deliver the accepted artifact to its named audience, retain outcomes and whole cost, and return after a material event. Treat the other 121 title nonmatches as not surfaced, not excluded; the five unresolved review positions and full lifecycle denominator remain null. 122 title-triaged · 1 GPT signal · 1 pinned-output bundle · 0 delivery/recheck chains
Mature delivery and afterlife without an AI role Exact-DOI discovery returns all 114 authoritative DOI records; 87 expose abstracts. AIDSVu is the sole abstract with explicit visualization, duration, observed-use, decision, and audience cues. Its paper reports ten years of public delivery, 501,527 unique users in 2019, named planning applications, and recurring governance. A 2026 data release and current FAQ add a later event and annual update/limitation contract. The held surfaces expose no AI visualization role, immutable artifact/build, version-bound measured audience or decision outcome, calibrated trust, affected-audience recheck, or whole cost. Aggregate reach and self-reported use examples cannot be borrowed into an AI feature. The 86 GenAI cue nonmatches are not exclusions; 27 DOI rows lack abstracts and eight stable keys lack DOIs. Keep AIDSVu as a non-AI comparator for durable public-data custody, audience surfaces, governance, and afterlife. For an AI claim, bind one immutable model/configuration and accepted artifact to intended-audience delivery, measured consequence, later material change, and affected-audience return. 114 DOI records · 87 abstracts · 1 non-AI afterlife · 0 complete AI lifecycles
Primary human evidence below the abstract layer All 27 no-abstract DOI rows have now been investigated. Six have full primary texts, 20 have primary abstract or official content, and one remains content-unassessed. The held cases separate controlled audience measurement, public-sector co-design/demo/bounded use, enterprise production intent, expert redesign, practice context, risk-ranking decision support, and a public non-AI delivery afterlife. The explicit GenAI and AI-visualization-role screens are empty for the 26 substantive rows. No row joins an accepted field release to a consequential affected-audience outcome and later recheck. B91 remains access-blocked and unassessed; task evidence, method description, co-design, demo exposure, production intent, expert review, practice application, and public reachability belong to different artifacts and stages. Pause generic B91 retrieval after three bounded passes and reopen only on an exact publisher abstract/body, accepted manuscript, correction, or new lawful deposit. Otherwise pursue the full B92 paper or one exact AI-assisted delivery-to-recheck episode. 27 investigated · 6 full · 20 abstract/official · 1 access-blocked/unassessed · 0 generic B91 targets
Authority repair below the DOI layer The eight stable non-DOI keys now yield three exact DOI repairs, two publication-year corrections, and one rejected foreign PMID. Custody separates four full texts, two primary abstracts, one issue excerpt, and one gated primary. The EMR cancer diary supplies the strongest practice receipt: increasing system-log use plus an independently recovered 11-clinician QUIS result with median satisfaction 4.38. Seven held content screens expose no explicit GenAI cue; gated B81 remains unassessed. The cancer diary exposes no AI role, immutable accepted build, log denominator, patient or decision outcome, later event, or affected-clinician recheck. The complete-review denominator remains null. Apply repaired identifiers without rewriting A208. Preserve theory, prototype, simulation, technical experiment, embedded use, accepted release, outcome, and recheck as separate stages. Recover B81, B112’s full paper/log denominator, partial B19/B142 text, or one of the 14 content-unassessed DOI rows. 8 keys · 3 DOI repairs · 1 operational comparator · 0 AI lifecycles
Human evaluation in the remaining DOI layer Nine high-signal rows were investigated; eight now expose a primary abstract or official project record and seven explicitly report participants or evaluators. InfoViP is the delivery near-miss: seven FDA safety evaluators supplied requirements, evaluated the prototype, and had suggestions addressed. An official FDA project page describes NLP and unsupervised learning and says an enhanced version will be installed in production. “Will be installed” does not confirm installation, acceptance, routine use, regulatory outcome, later change, or evaluator return. The eight content surfaces are not full papers. A subsequent residual pass reduces the content-unassessed DOI remainder from 14 to three. None of the eight contains an explicit GenAI cue. Seek an installation, acceptance, use, regulatory-impact, later-event, or evaluator-return receipt for InfoViP; recover full methods for the evaluated rows; and preserve evaluated, planned, installed, accepted, used, outcome-bearing, and rechecked as separate states. 9 investigated · 8 content surfaces · 7 human evaluations · 0 lifecycles
AI production component, QA afterlife, and operating-authority boundary A six-month internal real-work InfoViP evaluation records 20 unique reviewers and 58 submissions. The final CIOMS report says the component was approved for historical/live ETL, installed in AWS, integrated with AERS, and by 2025-07-30 had screened >30 million historical plus ~8,000 daily submissions. Case-series use is fully implemented and administrators monitor the pipeline/output. The now-closed Elsa/API/UI opportunity names Joshua Xu, Leihong Wu, and Oanh Dang as research mentors. Closure supplies no selection, onboarding, work, acceptance, or release receipt. The opportunity says a fellow is a nonemployee barred from inherently governmental functions. Official profiles support bounded research and project-lead roles, not InfoViP service operation, maintenance, QA execution, release approval, or authorization. An adjacent AI pharmacovigilance QA project describes a literature review and planned report but does not name InfoViP or a completed audit. The current HHS inventory still records implementation N/A and no ATO, and the Elsa 4.0/HALO launch never names InfoViP. Route research questions to the three mentors, but keep operator, maintainer, QA executor, release approver, authorization authority, immutable build, completed audit, application join, users/outcomes, return, and whole cost as separate required receipts. Treat opportunity closure as an administrative state, not execution. 3 research mentors · 0 operating authorities · opportunity closed · no InfoViP-Elsa join · 15 named · 0 full
Residual DOI primary recovery All 14 residual rows were investigated; eleven gained substantive primary content, including one full chapter that connects three UX experts and 25 unique problems to an implemented third version. A corridor stakeholder case and university-network case add practice context. At that pass, B92, B91, and B110 were metadata-only. No row exposed an explicit GenAI visualization role, accepted field release, routine-use denominator, consequential affected-audience outcome, later material event, or affected-actor recheck. Display context, expert review, stakeholder application, and multi-unit evaluation are not deployment. Preserve this historical 14 → 11 + 3 receipt, then apply the later B110 recovery without rewriting the earlier batch. Keep every receipt on its own artifact lineage. 14 investigated · 11 substantive · 1 full text · 3 unassessed at that pass
Final-three content and delivery recheck The exact three rows were rechecked. B110 now has an official bibliographic abstract, an immutable framework state, a pinned DiscoverWater application, and a same-named KU interface that is reachable in August 2026. At that pass B92 and B91 remained content-unassessed. B110’s live bytes are not bound to a SHA, and no receipt shows acceptance, continuous or ordinary use, consequential stakeholder outcome, accessibility acceptance, affected-audience return, or whole cost. Its held surfaces expose no AI visualization role. Treat source custody, a current public route, accepted build, ordinary use, audience outcome, and recheck as separate states. Bind live bytes to a release and measure named-audience use before making an adoption or impact claim. 3 rechecked · 1 newly substantive · 1 live surface · 2 unassessed at that pass
Live-build lineage and audience receipt audit A deterministic reconstruction proves that the dated live DiscoverWater page does not match the sole published v1.2 page state; the linked v2.0 source is R/Shiny. Partial data lineage remains: eight of twelve live dependency basenames occur in the pinned application tree, and two of three sampled asset pairs are canonically equal. An official 2018 record adds a limited prototype demonstration. No public manifest binds the live page to an immutable build. Analytics and comment hooks supply instrumentation, not traffic or audience evidence. Acceptance, ordinary use, consequential outcome, accessibility acceptance, later affected-audience return, whole cost, and an AI role remain missing. Preserve the dated page digest, source reconstruction, dependency coverage, and audience receipts separately. Reopen on an exact deployment manifest or version-bound audience record—not another same-name source link. 1 live page · 0 exact builds · 8/12 names · 2/3 canonical samples · 0 audience outcomes
Final-pair lawful primary recovery B92 now has an exact publisher abstract describing a multicriteria hydrogen-pipeline risk model, Monte Carlo simulation, Kendall’s tau rank comparison, and graphs for ranking sections and targeting mitigation. Three bounded B91 passes cover six named lawful surfaces. The abstract names no human sample or AI visualization role and supplies no release, intended-audience delivery, observed use, consequential outcome, later recheck, or whole cost. The full B92 paper is unheld. B91 remains content-unassessed; paused access is not a negative or world-level absence. Reuse B92 only as a non-AI uncertainty-and-ranking method comparator. Reopen B91 only on exact new lawful custody; otherwise move effort to the full B92 paper or one exact AI-assisted delivery-to-recheck episode with whole cost. 1 publisher abstract · 1 access-blocked/unassessed · generic search paused · 0 AI lifecycles
International and non-English use Multilingual POLYCHARTQA now tests 22,606 charts and 26,151 QA pairs across ten languages and finds substantial English/non-English gaps, especially in visual-language transcription. MM-JudgeBench finds language-dependent accuracy and bias across 25 languages, including a chart-centric judge subset. Real non-English authoring and reading, culturally situated chart conventions, code-switching, locally used analytics environments, more low-resource languages, and reader outcomes. Both benchmarks translate English-centric source material rather than sampling ordinary local work. Human-authored multilingual tasks and delivered-reader studies sampled across languages, chart conventions, institutions, and devices, with source distributions reported separately. Partly answered
Skill-package effectiveness Current packages provide advice, versioned references, contracts, executable checks, or browser inspection. SciVisAgentSkills improved quality in all ten paired suite-by-agent comparisons across 108 scientific-visualization tasks, although one completion measure fell. Independent tests of popular generic, dashboard, accessibility, mobile, and explanatory-visualization packages; decay after model releases, negative transfer, reader outcomes, and total instruction cost. No-skill, prose-only, narrow specialist, retrieval, verifier, and browser-evidence ablations repeated across current models, harnesses, and environments. Partly answered [Q12] [Q17]
Context-sensitive routing Established task and environment typologies explain why newsroom explanation, governed BI, open exploration, science, education, and operations impose different obligations. Which context dimensions actually change the best generator, critic, evidence bundle, or human gate—and which can safely remain shared. A factorial evaluation that varies audience, purpose, stakes, data custody, interaction, and delivery while holding the task family stable. Open [Q11]
Interaction and delivered state DashArena replays model-authored trajectories and materially improves human agreement by showing interaction evidence. Dashboard2Code tests callbacks and hidden state across 180 fixed-desktop Plotly Dash applications and 450 interaction tasks. Real-user exploration, responsive and mobile layouts, animation, popups, keyboard paths, authenticated applications, permissions, exports, assistive technology, device variation, and post-deployment state. Multi-viewport browser and human studies against the actual delivery surface, including state transitions, failure recovery, and non-happy paths. Measured but bounded [Q15]
Maintenance and total cost Twelve cost fragments are held. ChartAgent adds a quality/tool-call proxy; DV-World adds interaction charge; Selective TTS matches a declared partial inference budget. VisCoder2 adds a same-endpoint retry topology: over 888 tasks, VisCoder2-32B and GPT-4.1 both finish at 732 execution passes after 584 versus 714 revisions. The released debug paths discard the comparable usage and latency telemetry needed to price that difference. Ledger v7 separates matched declared partial budget, conditional attempt topology, observed route-wide use, and accepted-artifact cost. Zero of twelve held fragments has both equivalent observed per-arm route-wide cost and a frozen accepted-artifact denominator. Execution pass is not human or production acceptance, and no same-task record also joins preparation, delivery, a later event, independent recovery, maintenance authority, accessible-reader use, decisions, or calibrated trust under a current direct-work comparator. Freeze one multidimensional cap and acceptance contract; retain cold start through synthesis, effective topology/runtime state, calls/retries, tokens/image units, evaluator/tool work, accelerator use, latency, charges, human minutes, and accepted, rejected, abandoned, quarantined, and no-output states for every arm. Report the retry survival process and declared budget, observed use, eligible, attempted, candidate, accepted, delivered, and reader-successful ratios separately. Partly answered; 0/12 accepted-cost comparisons [Q15]
Same-artifact lifecycle joins Three dashboard cases—Prism, OpenClaw, and KubeStellar—cross delivery into later events. MAIDR adds a study surface, version floor, and later changes; Graphy adds repeated co-design. A six-receipt audit tests those five lines with PM4Py-UCM and DV-World. Every required receipt appears somewhere, but zero of seven rows joins repeated representative use, immutable tested build, exact exposure model, versioned release, later material event, and representative or actor-separated post-change recheck. One exact lineage clearing all six receipts, then extending into independent recovery, maintenance authority, accessible use, decisions, calibrated trust, or whole cost. The current null is scoped to named paper, repository, release, issue, and follow-up surfaces through 15 August 2026. Publish immutable exposure coordinates and re-run representative users after a named change on the released version. Preserve partial cells and search surfaces; never combine stages from different cases into a synthetic lifecycle. 0/7 full release-validation joins
Production recovery states The three dashboard afterlives all cross delivery into a later event. KubeStellar and Prism reach accepted repair, corrected delivery, and maintainer recheck. Zero case has an affected-user or independent-operator post-correction recheck, transferred maintenance authority, or all twelve states. The first same-case affected-actor return receipt or observed exercise of second-person release or incident authority. OpenClaw issue 30/PR 31 and Prism issue 81 were rechecked through 16 August 2026 UTC. Preserve actor, artifact/version, route, event, diagnosis, repair, delivery, maintainer recheck, affected-actor recheck, prevention, contribution, and authority as separate receipts. Issue closure is not recovery; contribution is not authority. 3 afterlives · 2 restorations · 0 independent recovery · 0 authority transfer
Reproducibility and transfer Several papers release code, data, or supplementary materials; others retain task sets, executors, model versions, or judges. Strong scores often depend on a renderer, grammar, or evaluator that does not transfer automatically. Independent reruns, cross-renderer tests, stable public tasks, and calibration against readers or domain experts. Full task, executor, judge, model-version, and failure-trace releases followed by third-party reproduction. Open

This ledger changes the research agenda in two ways. First, it prevents a single new paper from being narrated as “the gap is solved”: the relevant row moves only as far as the evidence permits. Second, it makes research pursuits testable. A skill-package survey must eventually lead to an ablation, not a popularity ranking. A specialized-vision survey must identify a same-task, equal-budget critic comparison, not merely another chart-question-answering leaderboard.

Active refresh triggers

The following releases or studies would materially change one or more rows:

Sources inspected

Pinned repositories

Papers read in full for this update

Human-outcome studies added by the gap-ledger pass

Benchmark and training studies for the moving-baseline analysis

Reader and accessibility study added by the open-question pass

Foundational task and environment frameworks

Current first-party product documentation

Model and harness milestone records

Earlier evidence retained in the comparison

The underlying research packet also contains VisEval, MAST, Cleveland–McGill MLLM testing, Does It Run, VIS-Shepherd, Text2Vis, CoDA, RealChart2Code, equal-budget single-agent results, synthetic-persona studies, and LLM visualization literacy work. Those sources continue to control the claims about reader simulation, integrity versus readability, equal-budget coordination, and benchmark coverage. This update does not silently promote abstract-only captures from that packet to full-text evidence.

Update log

Read or download the Markdown source