The state of AI-assisted data visualization research
Status: research snapshot, evidence cut 2026-08-18. Recheck by 2026-11-16, or earlier after a material model, harness, benchmark, registry, or analytics- assistant release.
Short version: executive summary. Forward view: What would unlock the next capabilities in AI-assisted data visualization? turns the moving-baseline evidence into dated, scoreable forecasts and keeps model-driven and technique-driven gains separate. Assessment view: How close is AI-assisted data visualization to the ideal? defines a context-sensitive ideal, scores nine current capabilities on one fixed maturity scale, reconstructs the same scorecard from 2017 through August 2026, and translates the forecasts into explicit score movements.
This report is for researchers, product teams, designers, and engineers who need to understand what AI-assisted visualization systems can currently do, which techniques appear to improve them, and how those claims are measured. It is a research synthesis and mechanism comparison, not a product ranking or buyer’s guide. Tools, agent skills, and commercial BI assistants appear as examples of techniques in practice; their documentation does not establish comparative quality.
Eight open-source repositories were inspected statically at the pinned revisions listed below. Eight core system papers were read in full, four initial benchmark and training studies were read for the moving-baseline analysis, three current commercial product surfaces were checked against first-party documentation, seven first-party model and harness release records were used for the shared timeline, and earlier papers in the underlying research packet were re-read where they controlled a comparison. Three established visualization frameworks were added to ground the distinction between task goals, narrative explanation, dashboards, and exploratory tools. A companion deep reading adds 18 agent-skill files or families, broader skills benchmarks, and the paired SciVisAgentSkills study; its report keeps registry installations, repository popularity, and behavioral evidence separate. A second specialized-vision report separates chart QA, parsing, OCR/layout, grounding, quality or integrity critique, and verifier/repair roles rather than treating every visual model as an interchangeable judge. A third companion, Learning data visualization when AI can make the chart, separates assisted performance from retained skill, distinguishes productive difficulty from removable friction, and forecasts how human practice should be reallocated as implementation becomes cheaper. A fourth companion, What AI-assisted data visualization benchmarks actually measure, translates benchmark names into their inputs, outputs, source distributions, graders, human context, supported claims, and nonclaims. No external repository, package, installer, skill, model gateway, or untrusted script was run. Repository tests and included evaluation results are observations about what the source contains, not independently reproduced results.
Start here: what agentic data visualization is trying to do
Agentic data visualization is not one task. A system may help a person discover a pattern, explain a finding, monitor a changing situation, make a decision, or produce and revise a durable visual artifact. These purposes can occur in a sequence—exploration may eventually become a published explanation—but they do not have the same success condition.
Rather than invent a new top-level taxonomy, this report borrows the discipline of Brehmer and Munzner’s established visualization-task typology: describe why the work is undertaken, what data and outputs it acts on, and how encoding and interaction support it. This matters acutely for agents. Returning a technically valid chart answers part of “how”; it does not show that the system served the intended human purpose.
| Primary goal | Typical environment | What good support looks like | What the agent must not optimize away |
|---|---|---|---|
| Explain and present | journalism, public explanation, narrative graphics | evidence supports a specific account; annotation, sequence, surrounding prose, and visual emphasis work together; defined readers can follow it | authorial claim and context, counter-reading, integrity, accessibility, publication judgment |
| Explore and discover | analyst work, open-ended research, exploratory tools | transformed data and alternatives remain visible; the person can steer, branch, correct, and reverse course | ambiguity, provenance, direct control, and the ability to abandon a bad line of inquiry |
| Monitor and respond | operational dashboards, shared awareness surfaces | current state, thresholds, anomalies, and action state are legible and trustworthy | freshness, stable measures, permissions, escalation rules, and operational consequence |
| Compare and decide | strategic and analytical BI, decision support | authoritative measures, fair comparisons, alternatives, uncertainty, and tradeoffs support a consequential judgment | semantic custody, metric ownership, role access, and human accountability |
| Produce, revise, and reuse | authoring tools, coding agents, scientific or interactive applications | the artifact executes, remains editable, reproduces its computation, and behaves correctly in delivery | domain constraints, edit locality, browser state, specialized scientific meaning, and final acceptance |
These are overlapping goals, not product bins. Segel and Heer characterize narrative visualization by the balance between author-driven narrative flow and reader-driven discovery. Sarikaya and colleagues show that dashboards also vary by purpose, audience, interaction, and data semantics, and differ materially from exploratory visualization tools. A newsroom interactive may contain both guided explanation and exploration; a dashboard may monitor, analyze, teach, or communicate. The important step is to declare the sequence and apply the right evidence at each transition.
That produces a general requirement: every system should preserve the question, data, decisions, artifact, and evidence. A context contract should then select the appropriate authoring and evaluation method. A universal generator, critic, or acceptance score would erase differences the system needs to understand.
What is being compared
This landscape contains several kinds of thing that are easy to blur together. A paper can describe a tool, a repository can implement a paper, and a benchmark can be used to test several tools. They still make different kinds of claim.
| Kind of object | Plain-language meaning | Examples in this review | What inspecting it can establish |
|---|---|---|---|
| Research paper or system | A proposed way to generate or edit visualizations, usually paired with an experiment | Raiven, NL2Dashboard, nvAgent, DashChat, PlotGen, NL4DV-LLM, Data Formulator 2 | What method was proposed and how it performed under the paper’s tasks, models, baselines, and graders |
| Open-source tool or framework | Runnable code intended for people or applications to use | Data Formulator, Lumen, LIDA, Vizro | What the current implementation can do and what checks it contains; not necessarily whether it improves human outcomes |
| Agent skill or plugin package | Instructions, references, templates, and sometimes scripts that shape how a general coding agent works | Vizro’s end-to-end flow, OpenAI visualize-data, AntV skills, Markdown Viewer Vega skill, claude-skillz |
What workflow and verification the package requests; only a with/without behavioral test can show whether an agent follows it or benefits |
| Benchmark or evaluation harness | A fixed set of tasks plus execution and scoring rules | DashArena, VisEval inside nvAgent, MatPlotBench inside PlotGen | Comparative performance under that harness; not automatic proof of production reliability or reader benefit |
| Commercial BI assistant | A vendor feature operating inside an established analytics product | Power BI Copilot, Tableau Agent, Looker Conversational Analytics | The current product contract and native sources of grounding; public documentation is not a comparative accuracy study |
| Technique | A reusable mechanism that may appear in any of the objects above | typed intent, compilers, semantic grounding, browser replay, direct manipulation | A candidate design principle. Its value must still be tested in the environment where it will be used. |
The objects fit together roughly like this, but the short pipeline hides who may decide what:
question + data
|
v
authoring system or product <--- an agent skill can guide this work
|
v
structured plan, DSL, or code
|
v
renderer + browser ---> executed chart or interactive dashboard
|
+--- deterministic checks test values and behavior
+--- a benchmark supplies repeatable tasks and scoring
+--- a human study tests how people author, understand, or decide
A reference anatomy assigns information and authority
A bounded audit of six held primary authoring or analysis systems—Data Formulator 2, Raiven, MatPlotAgent, Text2Vis, ViviDoc, and interactive task decomposition—supports an eleven-stage reference anatomy. The stages are a custody checklist, not one required agent topology:
| Stage | Information and legitimate authority | Boundary |
|---|---|---|
| Intent | A person declares purpose, audience, requested artifact, stakes, and stopping condition. | A prompt is not source, editorial, or publication authority. |
| Data and semantic grounding | Owners and deterministic contracts bind named sources, fields, values or metadata, units, measures, permissions, and uncertainty. | Fluent field selection cannot redefine an authoritative measure. |
| Inspectable representation | Assumptions, transformations, plan, state, encodings, interactions, constraints, and history can be inspected and rejected. | Hidden reasoning is not a governed specification. |
| Generation | A declared model and context propose code, specification, content, or an answer. | A candidate is not execution, correctness, or permission to publish. |
| Execution | A compiler, runtime, or browser turns exact code or a DSL into observable behavior and pixels. | A successful run does not prove semantic fidelity or reader value. |
| Critique | Deterministic checks, specialist models, and humans test named and different failure classes. | A critic is a bounded sensor, not a universal quality or acceptance authority. |
| Repair and history | Defects, feedback, changed state, prior state, non-regression results, and cost remain traceable. | A proposed repair is not a verified fix. |
| Human control | A named person can inspect, edit, branch, stop, reject, or escalate at consequential boundaries. | The opportunity to intervene is not a recorded acceptance decision. |
| Accepted delivery | An exact build, frozen rule, named approver, deployment, and delivery receipt travel together. | Benchmark pass, satisfaction, export, and local render do not qualify. |
| Reader or decision outcome | Intended people encounter the delivered surface under declared device, access, and task conditions. | Author ratings and model judges cannot stand in for readers. |
| Later maintenance | A later data, model, requirement, browser, or ownership event is repaired and rechecked with authority and whole cost retained. | Initial success does not establish updateability or handoff. |
The six-system audit finds 6/6 generation, execution, and some critique route; 4/6 meaningful human-control gates; and 0/6 exact-configuration accepted delivery, intended-reader outcome, or later maintenance. The 6/6 critique count is not architectural convergence: typed validation, runtime errors, visual-model feedback, and human inspection see different evidence and hold different authority. VisJudge can add one specialist critique signal; DashboardQA can test downstream agent interaction. Neither closes the human gates.
Primary anchors: Data Formulator 2, Raiven, MatPlotAgent, Text2Vis, ViviDoc, and interactive task decomposition.
Public acquisition is not adoption
The same custody rule applies to agent-skill popularity. Twelve named data-visualization registry listings carry 37,720 listing-summed install signals at the 14 August evidence cut. They collapse to ten parent repositories and eleven documented provenance lineages. The sum is not a count of unique people, teams, successful installations, invocations, or outcomes.
The registry calls its field a total deduplicated install count but does not publish the deduplication key or interval. At a pinned official CLI commit, an install event is assembled from selected skills after target results are collected; telemetry may be disabled or absent, and its schema contains no skill-invocation, task-outcome, retention, or organizational-acceptance event. The private audit therefore finds 12/12 public install signals and 0/12 held rows with observed real-task invocation, repeat use, organizational acceptance, or outcome/afterlife. Those zeros are bounded to the named public surfaces; private use remains unknown.
| Stage | Receipt needed | Named public rows |
|---|---|---|
| Listed | stable listing and captured package | 12/12 |
| Install signal | dated registry definition and count | 12/12 |
| Successful presence | exact package and version present after installation | 0/12 observed |
| Real-task invocation | task, model, harness, package, and execution receipt | 0/12 observed |
| Repeat use or retention | same actor or environment returns, with a denominator | 0/12 observed |
| Organizational acceptance | named owner, governed scope, and approval | 0/12 observed |
| Outcome and afterlife | accepted work, comparison, cost, later event, and recheck | 0/12 observed |
Use the skill-package report for the package-level table and evidence boundary. Use registry counts to choose what to inspect, not to make an adoption, procurement, or effectiveness claim.
Available controls are not governed deployment
The organizational-governance evidence also separates into two rails. Power BI/Fabric, Tableau, and Looker document meaningful enablement, data-boundary, and monitoring controls. Provider stories separately report named-feature use of Looker Conversational Analytics at Google Cloud Support and Copilot in Power BI among KPMG developers. Two longitudinal studies add adjacent governance- process evidence. Across those seven held rows, zero joins the complete same-deployment record.
| Evidence rail | Held rows | What it establishes | What it does not establish |
|---|---|---|---|
| Provider control contract | 3 | controls an organization can configure | the organization’s effective configuration, approval, use, or outcome |
| Named-feature organizational use | 2 | attributable provider-published reports of routine use | exact controls, audit and incident disposition, independent acceptance, or later recheck |
| Adjacent governance process | 2 | longitudinal authorization, evidence, ownership, or audit practice around organizational AI | an AI-assisted visualization deployment using the named provider feature |
| Complete governed deployment | 0 | — | no held row clears all nine receipts |
The join requires one feature and version, organizational authorization, enablement and scope, data or semantic authority, monitoring and audit, routine use, incident or exception disposition, an independently accepted outcome, and a later recheck. Provider controls from one row cannot be spliced to outcomes from another. This is a bounded public-surface result, not evidence that no organization holds a stronger private record. The practitioner account contains the provider-control, organizational-use, and source boundaries.
The eight core papers are not eight versions of the same product. Additional benchmark and training studies are introduced later where they answer the separate question of how the model baseline is changing.
| Paper | What it is, in plain language | Why it is in this comparison |
|---|---|---|
| Data Formulator 2 | A study of a chart-authoring interface where a person directly chooses visual encodings and asks the model for missing data transformations. | It is the clearest small study of mixed-initiative authoring, visible intermediate data, branching, and verification behavior. |
| Raiven | A scientific-visualization prototype where the model writes a restricted visualization language and a compiler produces linked 2D, 3D, and table views. | It supplies unusually strong bounded evidence that a DSL and compiler can eliminate many code-generation failures when the desired view is already specified. |
| NL2Dashboard | A dashboard architecture where the model produces a compact structured plan and uses small edit operations instead of rewriting the whole application. | It tests whether an intermediate representation improves controllability, edit locality, and token use. |
| nvAgent | A natural-language-to-database-visualization pipeline that prepares a schema, composes a visualization query, executes it, and repairs failures. | It provides a large benchmark and useful ablations for structured composition and execution-guided validation, especially across multiple tables. |
| DashChat | A conversational industrial-dashboard mockup tool with a dashboard DSL, design-pattern retrieval, structured edits, and history. | It combines a held-out prompt evaluation with a small user study of rapid prototyping and human correction. |
| NL4DV-LLM | A method that turns a natural-language question into an inspectable analytic specification: selected fields, analytical task, and candidate visualizations. | It shows why preserving explicit mappings and multiple interpretations can be more useful than jumping directly from prose to chart code. |
| PlotGen | A scientific plotting pipeline with separate numeric, text, and visual feedback passes around generated Matplotlib code. | It is evidence for multimodal feedback, while also illustrating why extra agents and extra model calls need an equal-budget causal test. |
| DashArena | A benchmark, browser executor, and human-calibrated judge for interactive dashboards—not a dashboard-authoring product. | It adds a serious measure of task-grounded analytical and interaction quality by replaying the author’s intended interaction sequence. |
Terms used throughout the report:
- Intermediate representation (IR): a structured, inspectable plan between a request and renderer code—for example fields, transforms, encodings, layout, and interactions stored as JSON-like data.
- Domain-specific language (DSL): a restricted language for one problem domain. It gives up some freedom so a compiler can reject or interpret output more reliably.
- Semantic model or semantic layer: a governed description of fields, measures, joins, permissions, and business meaning that sits above raw tables.
- Renderer: the software that turns a specification or code into pixels and interactive controls, such as Vega-Lite, ECharts, Matplotlib, or a BI product.
- Critic: a model pass that inspects code or a rendered artifact and proposes faults or repairs. A critic is another evidence channel, not an acceptance authority.
- Interaction trajectory: a recorded sequence such as open a tab, change a filter, and inspect the resulting state. A browser can replay it.
- Held-out task: an evaluation task not used as an example while building or prompting the system.
- Ablation: a controlled comparison that removes or changes one component to test whether that component actually caused the observed gain.
State of the field
This synthesis distinguishes repeated findings from environment-specific choices, direct negative or null evidence from merely unsupported claims, and demonstrated capability from open research questions.
Findings in one page
The field is not converging on a single agent architecture. It is converging on a more useful engineering pattern: reduce the part the model must improvise, externalize its decisions, and test the resulting artifact at the layer where failure can occur.
The strongest current techniques are:
- A typed visualization or dashboard representation between intent and rendering. Raiven, NL2Dashboard, ViviDoc, nvAgent, NL4DV-LLM, current Data Formulator/Flint, DashChat, and the best library-specific skills all constrain generation through a DSL, intermediate representation, or declarative spec. This improves syntax, edit locality, token efficiency, and the ability to validate individual decisions. It does not establish that the underlying question or takeaway is worthwhile.
- Deterministic work for deterministic claims. Compilers, schema checks, aggregation checks, executable code, browser traversal, replayed interactions, and regression comparisons outperform asking a model to pronounce an artifact correct. The model remains useful for ambiguity, semantic mapping, critique, and repair; it should not substitute for arithmetic or runtime evidence.
- Grounding in the actual data contract. Governed BI products increasingly bind natural language to semantic models, field descriptions, verified queries, permissions, sample values, and business glossaries. This is more consequential than assigning a generic agent an analyst persona.
- Mixed-initiative authoring. Data Formulator 2 and DashChat make natural language one control surface among direct manipulation, structured selection, visible transformed tables, history, branching, and reversion. This has stronger human-use evidence than prompt-only generation.
- Evaluation over the rendered artifact and intended use. DashArena’s new contribution is not another chart score. A system authors an interaction trajectory; a browser replays it; the judge sees task, screenshots, schema, and execution evidence. Human agreement improves materially when interaction evidence is present. This is the closest current answer to the earlier gap around communicative and analytical value, but it is not a measure of reader comprehension, learning, retention, or decision quality.
“Multi-agent” is therefore not a transferable technique by itself. It earns a place only when a role has a different information boundary or tool contract: schema access, transformation execution, visual inspection, browser interaction, or independent acceptance evidence. A planner, composer, and validator can help because they manipulate different representations and evidence, not because their labels simulate a team. Equal-budget single-agent results and several ablation findings still argue against agent count as a default quality lever.
Where the literature converges
The convergence is a common engineering posture, not a standard product architecture:
- Make intent inspectable. Independent systems repeatedly put a visible structure between the request and the render: an analytic specification, dashboard IR, DSL, semantic model, or chart contract. The vocabulary differs; the function is to expose fields, transformations, encodings, interactions, and edits for validation and correction.
- Use deterministic execution for deterministic claims. Recompute values, execute queries, compile specifications, inspect bindings, and replay browser actions. Models can map ambiguous language, criticize, and repair. They should not substitute for arithmetic or runtime evidence.
- Evaluate the delivered artifact. Returned code, structural similarity, and a nonblank render stop too early. Interactive work requires observable state and behavior; explanatory work requires final-context reading and a human gate when comprehension is claimed.
- Ground generation in the real data contract. Governed fields, measures, joins, permissions, verified queries, samples, units, and source vintage are more consequential than an “analyst” persona.
- Keep people able to steer and recover. The strongest authoring evidence favors visible transformed data, direct controls, small edits, history, branches, and reversion over prompt-only regeneration.
Where approaches legitimately diverge
Several competing approaches can each be right in different environments:
- Direct code versus a DSL and compiler. A restricted language can remove coding failures in a stable, specialized grammar; Raiven’s benefit is strongest in scientific visualization. Direct frontier models approach its reported performance in ordinary information visualization, where compiler overhead and an expressiveness ceiling may not pay for themselves.
- Prompt-first versus mixed initiative. Prompt-first creation is proportionate for a simple, reversible one-off. Exploratory or governed work benefits more from direct controls, visible state, history, and correction.
- Open-file tools versus governed BI. Portable code can work across renderers. A native BI assistant can instead inherit maintained semantic models, permissions, and worksheet state. Neither advantage transfers automatically to the other setting.
- One model versus specialized roles. A separate role can help when it owns different tools or independent evidence: schema access, execution, rendered inspection, or browser replay. Dividing one context into persona-labeled roles adds cost without establishing a new capability.
- Static inspection versus interaction trajectories. A static output may be adequately tested through data and final-render evidence. An interactive output needs action coverage and resulting state. A model-authored trajectory is still only a reproducible claim about intended use, not observed user behavior.
- Universal representation versus renderer-specific contracts. A universal IR promises portability; a narrow local contract preserves more expressiveness and is cheaper to maintain. The right boundary is an empirical question.
Direct negative, null, and conditional findings
- DashArena’s deterministic rules agreed with humans only 42.4% of the time, versus 79.8% for its full evidence judge. Removing interaction evidence reduced agreement by 8.1 percentage points.
- No DashArena model exceeded 86% render success or 74% interaction replay. Among 30 cases that executed cleanly but failed overall, 21 contained semantic defects.
- NL2Dashboard found diminishing returns from repeated critic rounds and recommends zero or one round, supporting capped repair rather than indefinite deliberation.
- Removing nvAgent’s processor slightly improved GPT-4o’s average result while harming weaker and multi-table settings. The component is conditionally useful, not universally beneficial.
- Raiven’s advantage was concentrated in scientific visualization; the DSL did not produce the same margin in ordinary information visualization.
- PlotGen’s human review marked 40.5% of outputs completely accurate and another 24.5% somewhat accurate. A sophisticated multimodal loop still produced many materially imperfect plots.
- DashChat users identified mock-data and domain mismatch, acronyms, direct-edit limitations, and styling limits. Rapid prototyping did not erase the need for domain context and correction.
Unsupported shortcuts are a different category
An unsupported claim is not proof of harm. It is simply insufficient evidence for acceptance. The AntV saved “success” results did not include browser render outcomes. The prompt-only packages in the initial sample have no behavioral with/without proof. LIDA’s code-and-text evaluator does not inspect the rendered chart. A validator that fails open is not an acceptance mechanism. Stars, installs, code returned, a nonblank render, synthetic personas, or one vision-model score may be useful inputs or baselines; none establishes data fidelity, integrity, reader understanding, or decision quality.
The model and harness baseline is moving
Shared timeline: September 2023 through August 2026
The three lanes below align general model and harness releases with the changing scope of visualization systems and benchmarks. They are a shared clock, not a causal model: release dates do not prove that a platform milestone caused a paper’s result, and publication dates lag the work. Scores remain comparable only inside the named studies.
The web presentation places every milestone on one proportional September 2023–August 2026 axis. Marker position encodes the date; label width does not encode duration. On narrow screens, the same evidence becomes a vertical chronology inside each explicitly named lane.
| Period | General model and harness baseline | Visualization systems and techniques | Benchmark and evaluation frontier |
|---|---|---|---|
| Sep 2023 | GPT-4V makes image input broadly available, enabling general models to inspect visual artifacts. | The main public capability question is still dominated by producing or reproducing individual static charts. | Existing evaluation mostly stops at code, structure, execution, or static-image similarity. |
| May-Oct 2024 | GPT-4o adds native multimodality; Anthropic’s computer-use beta adds screen, cursor, click, and typing actions. | Data Formulator 2 and NL4DV-LLM make transformed data, direct controls, history, and analytic specifications visible. | Plot2Code tests plot-to-code reproduction; VisEval tests 2,524 natural-language visualization queries across 146 databases with heterogeneous checks. |
| Feb-Aug 2025 | Claude Code brings a terminal coding agent; OpenAI’s agent tools bundle web, file, computer use, and tracing; GPT-5 targets coding and agentic work. | PlotGen, DashChat, and nvAgent add multimodal feedback, dashboard-specific languages, schema planning, execution, and repair. | Text2Vis combines data, questions, answers, code, and annotated charts across 1,985 tasks and isolates targeted feedback from generic prompting. |
| Jan-Aug 2026 | Hosted shell, computer environments, persistent workspaces, and reusable skills become first-party platform primitives. | NL2Dashboard tests a compact editable dashboard plan; Raiven tests a restricted scientific language and deterministic compiler. | RealChart2Code adds real data and multi-turn refinement; CharTide compares newer general and specialized chart models; Dashboard2Code and DashArena add state and replay; Chartography and FinChart-Bench test difficult professional reading; ChartDiff and multi-chart PolyChartQA add cross-chart reasoning; Chart-MRAG adds chart-bearing documents and retrieval; multilingual POLYCHARTQA and MM-JudgeBench expose language gaps in readers and judges. |
By the August 2026 evidence cut, the research question is no longer only whether a model can emit code for a plausible static chart. It is increasingly whether a model plus its harness can preserve meaning, construct a multi-view interactive artifact, exercise it in a browser, and supply evidence that it supports the analytical task. Reliability has not kept pace with that expanding scope.
The baseline has improved materially over the past 36 months, but the published record does not provide a clean visualization-specific learning curve from August 2023 to August 2026. Benchmarks changed, later papers often reran only a subset of models, judges changed, and current systems combine models with different prompts, tools, and test-time budgets. The defensible conclusion is a direction and a set of measured slices—not a universal annual improvement rate.
A fixed scorecard makes the remaining distance explicit
The companion capability scorecard applies one six-level rubric to nine dimensions: task framing; data, semantics, and provenance; construction; interpretation; integrity critique; steering and repair; interaction, responsiveness, and accessibility; reader or decision outcomes; and production, governance, and maintenance. The scale runs from “not demonstrated” through “outcome-proven.” It does not average the dimensions.
On that rubric, the August 2026 field is usable in bounded contexts for data grounding, ordinary construction, interpretation, and mixed-initiative repair; repeatable on bounded tests for integrity critique and delivered interaction; and only demonstrated—not established—for reader outcomes and the full production lifecycle. The same rubric can be applied retrospectively without pretending that unlike benchmark percentages share a common numerical axis. The historical reconstruction shows broad acceleration after 2023, but no common-core dimension yet crossing into field-wide deliverability. Context rows replace the generic ideal with the actual success condition for explanation, exploration, monitoring, decisions, production, scientific work, and reader assistance.
What the closest comparable slices show
The web presentation uses paired points for the before-and-after values below. Each measure keeps its own labeled scale and exact endpoints; readers may compare the two points inside a row, but not horizontal position across metrics. There is deliberately no combined score or common raw axis.
Plot2Code asks a multimodal model to reconstruct a scientific chart as code. An anonymous 2026 preprint re-reported direct, single-pass results on the same Python/Matplotlib subset and normalized scores over the full test set. Between the September 2023 GPT-4V result and the June 2025 Gemini 2.5 Pro result:
| Measure | Sep. 2023 | Jun. 2025 | Change over 21 months |
|---|---|---|---|
| Code executes | 84.1% | 87.9% | +3.8 percentage points |
| Text in the recreated chart matches | 48.5% | 71.7% | +23.2 points |
| Rendered-chart quality | 5.45 / 10 | 7.65 / 10 | +2.20 |
The large movement was in visual and textual fidelity, not basic code execution. This is a useful comparison, but it is one preprint’s reconstruction of one chart-to-code benchmark, not a field-wide time series.
CharTide, a peer-reviewed ACL 2026 training study, supplies a second same-paper comparison of GPT-4o and GPT-5 on three chart-to-code benchmarks. GPT-5 improved ChartMimic high-level similarity from 87.7 to 94.7, Plot2Code text match from 52.6 to 61.9, and ChartX’s five-point score from 2.61 to 3.59. Execution moved much less. The pattern is again uneven: generation is becoming more faithful, while harder visual reasoning and semantic alignment still leave substantial headroom. CharTide also shows that a small chart-specialized model can match or exceed a larger general model, so raw frontier scale is not the only route to improvement.
DashArena shows what the 2026 frontier can attempt that earlier chart benchmarks barely measured: a current general model can produce a multi-view interactive dashboard and a replayable intended-use trajectory from an open-ended task. Its top model was competitive with the anonymized human baseline in aggregate preference. Yet no model exceeded 86% render success or 74% replay success, and execution-clean semantic failures remained common. Breadth has expanded faster than reliability.
Dashboard2Code makes another part of dashboard behavior measurable. Its 180 Plotly Dash dashboard-code pairs cover 20 visualization types and eight callback patterns, with 450 interaction tasks. The best reported configuration scored 79.4 overall and 64.2 on the most complex interaction level. DOM access raised one model’s successful UI exploration from 20.61% to 64.58%, showing the value of exposing executable structure rather than relying on pixels alone. The benchmark also exposes a particularly dangerous failure: an interface can respond while hidden state or a transformation is factually wrong. It uses a fixed 1920×1080 viewport and excludes animation and popups, so mobile, responsive, and animated behavior remain outside the result.
Chartography asks a different question: can a model read difficult charts used in professional work? Its 100 practitioner-authored tasks span 12 domain labels and were independently verified by three experts per task. Thirty model configurations were run twenty times per task. The best tested configuration reached 45.0% mean pass@1. Greater reasoning effort helped in eleven of twelve paired comparisons, but the median gain was only 4.5 percentage points and longer reasoning often elaborated an initial visual misread. Because the tasks were deliberately screened for difficulty, 45% is not a prevalence estimate over all professional charts. It is strong evidence that success on basic visual-literacy or chart-QA sets does not imply dependable professional reading.
Scaffolding still helps, but generic instruction is already losing value
The most informative ablations do not say “more prompting is better.” They say that scaffolding helps when it adds a missing representation, tool, or evidence channel.
- In Text2Vis, GPT-4o’s direct pass rate was 26%. Three examples left it at 26%; retrieval plus three examples reached 31%. One structured answer-and-code feedback round reached 42%. Adding visual feedback improved the visual subscores but slightly lowered the final pass rate to 41%. Targeted feedback helped; examples alone were null; another feedback modality was not monotonically better.
- In the original Plot2Code study, adding detailed conditional requirements generally increased resemblance but reduced execution. Gemini Pro’s pass rate fell from 68.2% to 55.3%; the authors found no clear advantage for chain of thought or Plan-and-Solve over the default prompt.
- In nvAgent, removing the processor slightly improved GPT-4o on average while hurting weaker and multi-table cases. This is exactly what a moving baseline should produce: scaffolding that compensates for one generation of model can become redundant or obstructive for another.
The current operational baseline is also no longer a bare model call. Coding harnesses increasingly arrive with planning, repository search, execution, browser inspection, and packaged visualization guidance. A no-skill model is still a useful diagnostic control, but it is not the realistic alternative to a application-specific extension. The decision baseline must be the current model plus the normal harness, its default tools, and any ambient skills or instructions.
A current evaluation-readiness audit shows why that distinction matters. The product-owned historical record now resolves 14 case definitions, 18 run directories, and 13 judge directories, but the current public evaluation surface has only the narrow editorial sentinel. Of the five contexts required for a shared baseline—editorial, exploratory, governed BI, scientific, and interactive—all five now have executable source packets. The exploratory packet uses the exact CC0 Palmer Penguins table to test a real pooled/within-species sign reversal, branch preservation, row accounting, and render receipts. Its widespread use makes it a known-failure sentinel, not held-out performance evidence. The governed-BI packet uses a first-party synthetic semantic model, four roles, six verified queries, four refusals, and a fanout incident that produces $16,830 against the authoritative $7,730. That makes permission and recovery receipts executable, not production security or reliability. The specialized packet uses official 2024 ACS B19013 estimates and margins of error. Iowa, Kansas, Montana, and Wyoming span only $192; their largest internal z-score is 0.149508 against a declared 1.645 threshold. An exact point-only chart can therefore be mechanically valid while its strict winner story is unsupported. The interactive packet uses two fictional Riverbend maintenance snapshots, five canonical state fields, ten exact states, and seven isolated trajectories. Its critical transition makes selected T007 ineligible and leaves only T012 visible; selection, detail, accessible summary, URL, and export must change together. It also contracts history, reset, keyboard, narrow delivery, and a prepared V2 change. Historical auditability and source-packet readiness are not baseline readiness. Do not compare architectures until the exact current model/harness manifest is frozen and bare, normal-harness, and one-added-layer arms can emit equivalent receipts.
The first manifest preflight locked the five packet hashes, local toolchain, available-skill hashes, task order, three arm roles, and one 18-field receipt schema. A credential-free runner now clears the mechanical runner gate: five source checks run before it stages 15 isolated fixture-arm attempts; model inputs are separated from evaluator-only truth; artifact, render, trace, and check files are hashed outside the arm; and receipt identities fail closed. The editorial case required an explicit redaction because its source file contains the answer key. A whole-packet copy would have changed the task.
The successor preflight does not call the comparison ready. Six conditions
still block the first call: immutable served-model identity; exportable server-
harness configuration; same-model bare access; an authorized capped budget;
one versioned Vizier mechanism; and named acceptance authority or explicit
not-run treatment. The runner and successor verifier record zero calls and no
spend. A passing plumbing self-test, observed gpt-5.6 label, and local
codex-cli 0.147.0 remain neither a model result nor a server-side reproduction
receipt.
The same finding changes different decisions without changing the evidence:
| Audience | Decision layer |
|---|---|
| Practitioners and editors | Show the estimate and interval; describe the four middle states as unresolved under the declared test, not as a winner ladder. |
| Product, engineering, and BI leaders | Treat estimate plus margin of error as a bound pair through generation, semantic models, exports, and revisions; refuse requests that erase the uncertainty. |
| Visualization, HCI, and AI researchers | Score source fidelity, uncertainty propagation, statistical decision, refusal usefulness, expert verdict, and reader outcome separately. |
| Educators and accessibility specialists | Keep values, intervals, universe, unit, and comparison status available without color or hover; comprehension remains a human outcome, not a render check. |
The interactive packet adds a second audience decision layer:
| Audience | Decision layer |
|---|---|
| Practitioners and editors | After a filter changes, inspect the headline, marks, table, detail, shared URL, download, back button, and reset; a correct chart beside stale detail is still broken. |
| Product, engineering, and BI leaders | Derive every surface from one canonical state, invalidate an ineligible selection atomically, and require replayable pilot receipts before treating a demo as application evidence. |
| Visualization, HCI, and AI researchers | Start each trajectory from a declared isolated state and score visual fidelity, behavior, responsive delivery, keyboard equivalence, maintenance, and human usefulness separately. |
| Educators and accessibility specialists | Require every filter, ticket, reset, and export without pointer input and announce state changes without stealing focus; actual assistive-technology use remains not run. |
The manifest preflight changes how those audiences should read comparisons:
| Audience | Decision layer |
|---|---|
| Practitioners and editors | Ask whether model, tools, context, token budget, retries, and reviewer access changed together before crediting one prompt or skill. |
| Product, engineering, and BI leaders | Require an executable manifest, equivalent arm receipts, and total cost before pilot or procurement claims; a model label plus CLI version is insufficient. |
| Visualization, HCI, and AI researchers | Report missing configuration before outcomes, keep a same-model bare diagnostic distinct from another local model, and preserve failures rather than imputing scores. |
| Educators and accessibility specialists | Keep expert and human fields structurally present as not-run; missing keyboard, assistive-technology, learning, or comprehension evidence cannot be simulated. |
The neutral runner turns the same answer into seven concrete decision products:
| Audience | Runner translation |
|---|---|
| Visualization and data practitioners | Inspect the actual model-input manifest; an answer key, expected value, or acceptance rubric in context is part of the intervention. |
| BI and analytics leaders | Require an exportable pilot directory with source, role, semantic context, artifact, renders, checks, failures, human state, and cost—not a screenshot plus score. |
| Visualization, HCI, and AI researchers | Stage before inference, preregister the visibility boundary, hash outputs outside the arm, and report missing receipts as missing. |
| Data journalists, graphics editors, and newsroom developers | Preserve the draft, source, renders, revision, and editorial judgment separately; do not let generation see the private answer key or let a summary erase failure. |
| Product, engineering, and tool teams | Put neutrality in the filesystem: allowlisted input, isolated attempt roots, external output inventory, and a separate versioned provider adapter. |
| Educators, data-literacy, and accessibility specialists | Deterministic checks can show that files or states exist; they cannot stand in for comprehension, transfer, keyboard experience, or assistive-technology use. |
| Executives, editors, and broad AI readers | The evaluation plumbing is testable without spending. No model has run, and six model, budget, intervention, and acceptance decisions remain. |
One layer means one controlled difference
A skill package is an installation and use surface, not automatically a causal unit. The two pinned public Vizier skills combine reader framing, form and pattern retrieval, honesty questions, encoding and palette guidance, deterministic commands, actual-artifact inspection, optional corpus retrieval, and critique composition. Testing either whole package would show the value of that bundle under its budget, not which mechanism mattered.
One candidate is now prepared but deliberately unselected: artifact-evidence
reconciliation checkpoint v0.1. It enters once after the first executable
artifact and neutral output inventory, before repair or finalization. For each
decision-bearing visible claim or state, it records the artifact hash, available
source/transform/query/calculation/branch/state evidence, and match,
mismatch, or unresolved. Missing states remain missing; discrepancies are
frozen before any bounded repair. Overall pass is valid only when every
surface matches, no required state is missing, and no repair remains. The one
failure target is silent divergence between the delivered artifact and the
evidence behind it.
Use this five-part test before calling anything “one layer”:
- one named mechanism and one expected failure target;
- one lifecycle insertion point and one versioned output receipt;
- identical model, harness, inputs, tools, budgets, repair cap, and evaluator;
- evaluator answers and sentinel mappings kept outside model input; and
- an explicit owner decision distinct from a passing protocol verifier.
The same candidate changes seven audience decisions without changing its evidence status:
| Audience | One-layer decision |
|---|---|
| Visualization and data practitioners | Ask which one checkpoint changed; do not credit a prompt when form, context, tools, retries, or critique changed too. |
| BI and analytics leaders | Name the one control, insertion, wrong number or scope leak it targets, proof directory, false alarms, and review cost. |
| Visualization, HCI, and AI researchers | Register the delta, exclusions, visibility boundary, common budget, and output schema; protocol readiness is not an effect estimate. |
| Data journalists, graphics editors, and newsroom developers | Call a whole workflow a bundle; attribute one editorial mechanism only when it alone changed and the failed draft remains inspectable. |
| Product, engineering, and tool teams | Version framing, retrieval, construction guidance, deterministic tools, artifact inspection, and critique composition as separable lifecycle hooks. |
| Educators, data-literacy, and accessibility specialists | An evidence mismatch check can find inconsistent labels or states; it cannot supply learning, comprehension, keyboard, or assistive-technology acceptance. |
| Executives, editors, and broad AI readers | One candidate is testable, not selected or proven. Owner ratification and the other five gates still precede comparison. |
A pass needs the right authority
“Human review” is not one interchangeable box. A context reviewer decides whether an artifact supports the stated job. A domain expert checks semantics and interpretations that deterministic invariants cannot settle. A knowledgeable accessibility evaluator assesses a declared conformance scope. Relevant disabled users show how a delivered task works on their setups. Intended users accept or reject the artifact for the decision. A later independent maintainer supplies handoff and change evidence.
These roles cannot be filled by an automated critic. The W3C evaluation overview says no tool alone determines accessibility. Its guidance on involving users also says user evaluation complements rather than replaces standards work and must report participant scope without overgeneralizing.
The prepared E0 candidate makes that boundary executable:
| Outcome lane | Eligible authority | Mechanical-only state |
|---|---|---|
| Contestable context review | Named graphics/data editor, analyst, finance owner, stakeholder, or operations manager appropriate to the fixture | not-run |
| Domain-expert review | Named source, semantic-model, survey-methods, measurement, or operations expert appropriate to the claim | not-run |
| Accessibility conformance | Named knowledgeable accessibility evaluator | not-run |
| Accessibility user evaluation | Relevant disabled users, with task, setup, and participant scope retained | not-run |
| Intended-user acceptance | Named person in the fixture’s declared audience | not-run |
| Maintenance handoff | Later independent maintainer working from the retained artifact and history | not-run |
Five fixtures times six lanes yields 30 explicit not-run records. A
scope receipt rejects a human pass without a named human authority and rejects
an automated tool as that authority. A deterministic result may therefore be
reported only as a mechanical pass for the named fixture and arm—not as
unqualified acceptance, accessibility, usefulness, comprehension, maintenance,
or Vizier efficacy. The research owner has not selected mechanical-only scope
or named people, so this is a candidate policy with zero model calls and six
remaining preflight blockers.
The audience consequence is concrete:
| Audience | Authority decision |
|---|---|
| Visualization and data practitioners | Keep checks, editors, domain reviewers, intended readers, and later maintainers as separate stopping conditions. |
| BI and analytics leaders | Name finance, semantic/governance, user, incident, and maintenance ownership before upgrading a synthetic mechanical result. |
| Visualization, HCI, and AI researchers | Register authority, independence, information access, participant scope, and evidence; treat not-run as missing outcome data, not zero or pass. |
| Data journalists, graphics editors, and newsroom developers | A palette collision can be computed; source binding, editorial acceptance, accessibility, and reader trust still need their own authorities. |
| Product, engineering, and tool teams | Encode authority as typed receipt state and reject models or tools that attempt to fill a human lane. |
| Educators, data-literacy, and accessibility specialists | Keep conformance evaluation and disabled-user experience distinct, and report the limits of each participant scope. |
| Executives, editors, and broad AI readers | Five fixtures have one verified authority candidate; 30 human states remain unrun, no scope is selected, and no person was contacted. |
Six answers, then one release
The six remaining E0 blockers are now one decision register, not one blanket approval. It offers 17 admissible paths across four authority roles and keeps each answer separate:
| Gate | Who can answer | Advancing answer |
|---|---|---|
| Served-model identity | Model provider or runner | Supply an immutable snapshot and served receipt. |
| Server-harness custody | Harness provider or runner | Export the reasoning/output envelope, harness build, instruction digests, and tool-service versions. |
| Bare-model interface | Provider for access; research owner for protocol | Supply a same-model raw surface, or amend the protocol to drop that arm and narrow the claims. |
| Authorized budget | Research owner | Name the provider/model route, price basis, call/token caps, and stop condition, including for an evidenced zero-cost route. |
| One Vizier layer | Vizier owner | Ratify the exact v0.1 candidate; revision or rejection keeps the layer unselected. |
| Acceptance authority | Research owner | Ratify mechanical-only scope or require a real named-authority protocol before running. |
An unavailable surface, declined budget, requested revision, or rejection is a valid recorded decision, but it cannot make the run ready. Even six advancing answers do not trigger a call: the research owner must separately release the hash of that resolved decision set against a named run plan, call cap, and stop condition. Decision receipts may identify a non-secret route and hash evidence; credential values never enter custody.
The verified candidate rejects wrong-role, incomplete, secret-bearing, partial, and six-answer-without-release states. That is protocol integrity, not approval. All six selections and the final release remain null, spend and readiness remain false, no person was contacted, and no model was called.
The audience translations make the same state useful without changing it:
| Audience | Decision-register translation |
|---|---|
| Visualization and data practitioners | Six blank answers mean “not run,” not “almost passed”; ask which model, harness, budget, layer, and acceptance scope were actually selected. |
| BI and analytics leaders | Use the register as a pilot approval sheet: provider identity, exportable envelope, cost cap, named control, acceptance policy, and accountable release. |
| Visualization, HCI, and AI researchers | Publish gate receipts, amendments, and the decision-set digest; preserve unavailable surfaces and missing human outcomes rather than imputing them. |
| Data journalists, graphics editors, and newsroom developers | Keep provider facts, editorial/product choice, reader acceptance, and publication judgment in separate hands. |
| Product, engineering, and tool teams | Implement typed fail-closed state; provider adapters return non-secret receipts and never mutate owner choices or infer release from credentials. |
| Educators, data-literacy, and accessibility specialists | Harness and budget approval do not establish learning, comprehension, conformance, or disabled-user experience; those lanes stay not-run until performed. |
| Executives, editors, and broad AI readers | The experiment is prepared but not authorized: six decisions and one release remain, with zero calls and no efficacy result. |
Score layers, not “quality”
The prepared E0 scorer has no overall score, winner, rank, universal pass, or efficacy field. It reports two mechanical tiers—computable checks and checkable obligations—plus six named nonmechanical lanes: context review, domain review, accessibility conformance, disabled-user evaluation, intended-user acceptance, and maintenance handoff. Cost stays beside those outcomes.
A mechanical tier passes only when it contains at least one check and all
checks pass. One failure yields fail; an empty tier or any not-run yields
incomplete. Each nonmechanical result retains its exact authority, rationale,
scope, and evidence. not-run remains missing data, and one human role cannot
sign another role’s lane.
The zero-call verifier created 15 synthetic fixture-arm ledgers and preserved
90 nonmechanical not-run records. It rejected unnamed and cross-lane human
authority and injected aggregate score/pass fields. That establishes scoring
plumbing, not 15 successful artifacts or a Vizier effect.
| Audience | Tier-separated decision |
|---|---|
| Visualization and data practitioners | Read exact checks, contextual review, intended-user acceptance, and maintenance separately; a clean render is not a good outcome by itself. |
| BI and analytics leaders | Require measure/permission checks, owner review, user acceptance, incident evidence, cost, and handoff as separate rows before adoption. |
| Visualization, HCI, and AI researchers | Compare arms by predeclared lane, preserve missingness and parse failures, and report trade-offs rather than a composite leaderboard. |
| Data journalists, graphics editors, and newsroom developers | Keep source/mark checks, editorial judgment, accessibility, reader evidence, and publication authority distinct. |
| Product, engineering, and tool teams | Emit typed receipts from each evaluator and reject cross-lane authority or schemas that add an overall pass. |
| Educators, data-literacy, and accessibility specialists | Conformance, disabled-user task evidence, learning, and transfer answer different questions and cannot substitute for one another. |
| Executives, editors, and broad AI readers | “15 ledgers verified” means the scorer works on synthetic records; no AI output or human outcome succeeded. |
Integrity and readability need a blinded test
R4 now has a preregistered pilot and five frozen inputs, not a result. Five opaque-ID SVGs, five neutral briefs, and nine source-evidence files make one 19-file model-input allowlist. Five context mappings and known-defect records stay in separately hashed evaluator-only custody. The model can inspect the same visual and evidence under either lens without receiving the answer key.
The two lenses now also have one hashed prompt and one fail-closed raw-response schema each. A response may complete with evidence-grounded findings, complete with no findings, or explicitly abstain. Twelve synthetic adverse records show that aggregate fields, mixed lenses, wrong identities, incoherent evidence, and findings attached to an abstention are rejected. This is tested local plumbing; it does not show that a provider supports the format or a model will follow it.
Three exact-model slots, the independent human data editor, calibrated tie rule, budget, release, and results are still empty. Future models produce separately blinded integrity and readability observations without scoring themselves; the editor later scores them against hidden defect truth and a declared readability rubric. Only the two within-lens partial orderings are compared. They never become one quality score.
The remaining approval path is now explicit. Five gate categories become seven attributable receipts: one exact served receipt for each of M1, M2, and M3; editor-signed role acceptance; same-editor calibration; a research-owner capped budget; and a final research-owner release over the exact six- prerequisite digest. The budget must cover 30 primary records—three models by five stimuli by two lenses—with retries capped separately.
Twenty advancing and non-advancing receipt paths validate. Sixteen synthetic adverse states show that wrong authority, cross-subject choices, incomplete or unsafe evidence, duplicate/partial prerequisites, a refusal, missing release, wrong digests or hashes, different editors, route mismatch, call-cap mismatch, and duplicate model identity remain blocked. This tests the handoff mechanism; all six real prerequisites are pending and release is false.
The editor’s future calibration task is now prepared but blank. It contains ten baselines—five stimuli under each lens—and ten synthetic response-quality vignettes, including two A/B-swapped repeats. The response can complete, request revision, or state that usable resolution cannot be established; fourteen adverse custody, coverage, repeat, state, and aggregate-result records fail. No editor, appointment, calibration, resolution, or tie release exists.
| Custody layer | What is present | What it licenses |
|---|---|---|
| Model input | 5 visuals + 5 briefs + 9 evidence files, each path and hash allowlisted | A future evidence-grounded audit after provider and owner release |
| Evaluator only | 5 defect mappings, expected findings, false-allegation guards, and 9 exact origin bindings | Later scoring by the named editor; never model context |
| Prompt/output | 2 lens prompts + 2 schemas; finding, empty, and abstaining states | Inspectable raw observations, not self-scores or a model comparison |
| Severe checks | Duplicate IDs, truth-path leakage, wrong hashes, and missing evidence are rejected | Confidence in staging boundaries, not model quality |
| Response checks | 12 aggregate, mixed-lens, identity, evidence, and state failures are rejected | Confidence in the response boundary, not provider compliance |
| Visual inspection | Corrected label overlap, an undeclared truncated baseline, and hidden export text before freeze | A cleaner pilot input set, not an audience or accessibility result |
| Receipt | Who supplies it | Current state |
|---|---|---|
| M1, M2, M3 served routes | Provider or runner for each exact slot | All three pending |
| Editor acceptance | The named independent data editor | Pending; no person contacted |
| Editor calibration | The same accepted editor, with retained evidence | Pending; resolution and tie behavior remain null |
| Capped 30-primary-call budget | Research owner | Pending; no spend authorized |
| Exact-set release | Research owner, naming the six-receipt digest and frozen hashes | False; no call authorized |
| Audience | R4 decision |
|---|---|
| Practitioners | Ask for a bounded observation, exact visual/source locator, and scope; an empty finding set or abstention is evidence, not a formatting failure. |
| BI leaders | Require completed integrity records to consult both the visual and source evidence; one fictional semantic model is still not a procurement score. |
| Researchers | Preserve raw lens records, empty findings, abstentions, and parse failures; apply one human ordering later instead of collecting model self-scores. |
| Newsrooms | Models produce cited observations only. A named data editor retains hidden-truth, false-allegation, score, tie, and publication authority. |
| Product/tool teams | Validate returned JSON against the frozen lens schema and exact stimulus allowlist; record provider support in a served receipt rather than assuming it. |
| Educators/accessibility specialists | Readability observations are not comprehension, learning, conformance, or disabled-user results; integrity observations are not reader performance. |
| Broad AI readers | Two prompts and two schemas pass synthetic tests; zero models or people have answered them, so no winner or divergence exists. |
The same register changes the decision question by audience. Practitioners can ask whether all seven receipts are inspectable. Analytics leaders can use them as a pilot-control sheet. Researchers can retain refusals and provider failures as results about the protocol rather than silently changing it. Newsrooms keep provider, adjudication, research release, and publication authority separate. Product teams can re-hash a route adaptation instead of hiding it in the runner. Education and accessibility readers can see that editor calibration is not learner or disabled-user evidence. Broad readers get the shortest honest status: the approval path is testable, but every real answer is blank.
For practitioners and analytics leaders, the additional check is whether the editor wrote the baseline before seeing model output. Researchers should treat the swapped repeats as a consistency check, not inter-rater reliability. Newsrooms retain editorial/publication authority; product teams keep the packet evaluator-only; education and accessibility readers should not read review- scale calibration as comprehension or disabled-user evidence. Broad readers can say only that the editor now has a concrete future task.
What can and cannot be projected
The observed direction supports three bounded forecasts:
- Likely to commoditize: syntactically valid chart code, conventional chart selection, basic styling, routine repair, and generic “inspect your render” advice. These should be short, replaceable defaults, not a large permanent application-specific doctrine.
- Likely to migrate into models or harnesses: generic planning, self-review, browser/tool use, and library navigation. A visualization system should consume these when they work and avoid duplicating their orchestration.
- Unlikely to be solved by baseline capability alone: the local question, audience, semantic definitions, authoritative data, denominator, source custody, consequence, publication boundary, environment-specific evidence, and real-reader outcomes. These are facts and authorities the model does not acquire merely by becoming better at code or vision.
Linear projection would be false precision. The practical forecast is that the value of generic instruction will decay fastest, the value of executable local evidence will persist, and the value of context and authority boundaries will increase as agents become capable of taking more consequential action.
How evidence was graded
The same word—“works”—covers incompatible outcomes in this literature. This report keeps seven layers separate:
| Layer | What can be established | Typical evidence | What it does not establish |
|---|---|---|---|
| Structure | Required fields or files exist | schema validation, AST checks, spec comparison | execution, correct values, usefulness |
| Execution | Code compiles/runs and a nonblank artifact appears | sandbox, renderer, browser, console/network checks | correct binding or interpretation |
| Data fidelity | Values, aggregations, filters, and joins match source data | deterministic recomputation, exact comparisons, query receipts | legibility or audience value |
| Visual integrity | Encodings are coherent and nondeceptive | mark/encoding checks, render inspection, targeted critic | ease of reading or insight |
| Interaction | Controls and linked views behave as intended | browser replay, action coverage, before/after state | that the interaction supports a good analysis |
| Analytical support | The artifact helps address a task | task-grounded comparison, expert study, realistic workflow | learning, retention, or downstream decision quality |
| Reader outcome | A defined audience understands, remembers, or decides better | controlled human study with audience/task measures | generalization beyond that population and context |
Passing a lower layer is a prerequisite, not a proxy for all higher layers. “Rendered,” “valid,” “high similarity,” and “preferred by a vision model” are different claims.
Evidence labels used below:
- A — comparative human/outcome evidence: a relevant controlled user or human-calibrated task study, with material caveats still reported.
- B — comparative benchmark evidence: a held-out or public task set with explicit metrics and baselines, but not a direct reader outcome.
- C — executable mechanism evidence: inspectable implementation, tests, or evaluation fixtures without persuasive independent efficacy evidence.
- D — product or author claim: first-party description, demo, popularity, or anecdote without a usable comparative outcome.
Technique map
| Technique | Representative systems | Best current support | Environment fit | Failure boundary | Current evidence-based posture |
|---|---|---|---|---|---|
| Prompt-only visualization guidance | claude-skillz data-visualization, Markdown Viewer Vega skill |
C/D: inspectable instructions; no behavioral comparison found | one-off, low-risk chart creation where the renderer is already known | prose can be stale, ignored, or internally inconsistent; no data or render proof | Reject as a sufficient workflow; retain only as baseline |
| Versioned library constraints plus retrieval | AntV G2/G6/X6 skills | C: large reference corpus, retrieval datasets, structural checks, included generation results | coding against a fast-changing visualization API | reported “success” can mean response returned; saved results were not browser-rendered | Try as a retrieval intervention, not as quality proof |
| Chart contract before implementation | OpenAI visualize-data; Vizro design specs |
C: explicit contracts, runtime tests, required specs and receipts | coding agents, reports, dashboards, multi-surface delivery | instruction compliance is probabilistic; contract can document a wrong question | Borrow now |
| Structured analytic or visual IR | NL4DV-LLM, nvAgent VQL, NL2Dashboard IR, RaivenDSL, ViviDoc SRTC, Flint | A/B in bounded tasks plus ViviDoc’s same-team C-grade ablation | repeated generation/editing, scientific views, dashboards, interactive documents, renderer portability | schema can exclude useful forms; semantic binding can still be wrong; creator satisfaction is not reader outcome | Borrow the principle; test the smallest local IR |
| Deterministic compiler/renderer | Raiven, NL2Dashboard, Flint, Vega-Lite-based tools | Raiven B/A in fully specified reproduction; NL2Dashboard B | scientific visualization, stable dashboard grammar, controlled environments | does not discover the analytical question; expressiveness ceiling and compiler defects remain | Try against direct code at equal task/budget |
| Direct manipulation plus natural language | Data Formulator 2, DashChat, Tableau Agent | Data Formulator 2 and DashChat A within small studies | exploratory analysis and prototype negotiation | user studies are small; current products have evolved beyond published evaluations | Borrow interaction pattern |
| Persistent history, branching, reversion | Data Formulator/Data Threads, DashChat history | A/C: observed in studies and current implementation | nonlinear analysis and stakeholder iteration | provenance can record bad branches without detecting them | Borrow now |
| Semantic-layer grounding | Looker/LookML, Power BI semantic models, Tableau field metadata; current Data Formulator connectors | C/D: strong mechanism and governance rationale; limited public comparative outcome evidence | governed enterprise BI with established models | semantic layer can be stale or wrong; unmodeled questions remain hard | Borrow as an input contract when one exists; never infer semantic truth from availability |
| Execution-guided repair | nvAgent, PlotGen lexical loop, Lumen agents, Vizro testing | nvAgent B; other evidence mixed | syntax/API-heavy generation and multi-table queries | fixes what throws, not silent semantic error; repeated loops add cost | Borrow bounded repair after typed failure |
| Rendered-image critique | VisJudge, Raiven VMPC judge, PlotGen visual agent, NL2Dashboard critic | VisJudge B on its expert-adjudicated quality rubric; Raiven B with human correlation; other evidence weaker | visible composition, readability, mark, label, and bounded perceptual failures | a screenshot critic cannot verify source fidelity, interaction, responsive states, or reader outcomes | Try only as one evidence channel |
| Chart-specific perception and tools | ChartAgent, ChartREG++, chart parsers, OCR/layout models | Bounded specialist gains in numeric QA, chart-mark grounding, and document parsing; current general models lead some new transfer tests | exact extraction or localization failures that a general critic cannot resolve reliably | parsing is not critique; tools still fail; specialist benchmark fit can decay quickly | Route to a measured component failure, not to “quality” in general |
| Deterministic data/integrity checks | Vizro aggregation/color scripts; Raiven data-hallucination checks; local regression guards | B/C and strong causal fit | any generated chart with inspectable source data | check coverage is necessarily partial | Borrow now and expand by failure class |
| Browser and interaction evidence | Vizro Playwright workflow, DashArena executor | DashArena A/B; Vizro C | interactive dashboards and browser-delivered reports | scripted path may miss latent behavior; clean execution can hide data errors | Borrow now for interactive surfaces |
| Model-authored replayable analytical trajectory | DashArena | A/B: 234 tasks, human-calibrated pairwise judging | open-ended interactive dashboard generation | trajectory is partial and model-authored; benchmark is Tableau-seed-biased | Try as an evaluation receipt, not as product behavior |
| Multiple role agents | LIDA, PlotGen, nvAgent, Lumen, DashChat, NL2Dashboard | mixed: nvAgent/DashChat positive, equal-budget general evidence negative, several ablations nuanced | heterogeneous access/tool boundaries or parallel independent work | role theater, correlated critics, token/latency growth, weak baselines | Use only where the role owns a distinct capability or evidence boundary |
| Synthetic reader/persona panels | LIDA goal personas; broader synthetic-user literature | weak for audience validity; negative evidence on faithful simulation | brainstorming possible questions only | false confidence about real readers, difficulty, aesthetics, and persuasion | Reject for acceptance or audience claims |
Research and implementation landscape
This landscape shows where research techniques appear in working systems, agent-skill packages, and governed BI products. The entries are examples and evidence records, not product scores: a paper can support a bounded performance claim, a repository can expose an implementation, and product documentation can describe a contract, but those are different kinds of evidence.
Full tools and research systems
| Offering | Actual mechanism | Evidence read | Assessment |
|---|---|---|---|
| Data Formulator | Natural language handles transformations and agentic exploration while a GUI handles explicit visual encodings; visible tables, code/explanations, threads, branches, and current Flint semantic specs support refinement. | Data Formulator 2: eight participants reproduced 16 charts and 12 nontrivial transformations; all completed the tasks, with distinct branch/depth strategies. Current 0.8 alpha is substantially broader than the studied 2024 prototype. | A/C. Strongest authoring interaction pattern. The user study supports learnability and verification behavior, not long-term analytical correctness or the efficacy of the 2026 agent stack. |
| Raiven | The model sees metadata and emits RaivenDSL; a deterministic compiler creates coordinated 2D, 3D, and tabular views with controls. | 100 fully specified prompts; 100% compile and .988 VMPC versus .800–.867 compile and .678–.721 VMPC for direct-code baselines. Seven visualization experts completed three replication tasks; five preferred Raiven and two had no preference. | A/B, bounded. Compelling for unfamiliar scientific/3D prototyping. Most gains are in SciVis; InfoVis baselines nearly match. Tasks reproduce specified views rather than discover questions. Authors were the graders, and the tool can still encode logically misleading order/null choices. |
| NL2Dashboard | A compact IR separates analysis/content/layout from deterministic rendering; atomic modification operators avoid rewriting a dashboard; planner, coder, and optional visual critic operate around the IR. | Ten tables across finance, education, and government; seven modification classes. It scored 11.89/15 in generation and 11.93/15 in modification under an LLM judge, completed all edit tasks, and used much lower output-token/dashboard ratios than web-product baselines. | B, promising but not decisive. The IR/edit result is strong. The comparison uses different model interfaces/backbones, only ten source tables, an LLM quality judge, no user study, and no equal-budget control. The paper reports diminishing critic returns and recommends zero or one round. |
| ViviDoc | A DocSpec decomposes interactive units into State, Render, Transition, and Constraint before code; authors can edit the plan, style, and result. | 101 topics across 11 domains; a same-pipeline ablation across three backbones; 36 blind-rated outputs; and 12 people authoring two documents each. The largest reported interaction-quality lift from DocSpec was 41%. | C, useful authoring mechanism. Same-team evidence and partly model-judged metrics. DOM change and creator satisfaction do not establish source truth, reader learning, accessible or production delivery, maintenance, or autonomy. |
| nvAgent | Processor filters/augments database schema, composer uses sketch-and-fill to make VQL, validator translates to Python and iterates on execution errors. | ACL 2025; VisEval has 2,524 NL/visualization pairs over 146 databases. With GPT-4o it reached 85.63% single-table and 81.07% multi-table pass rate, +7.88 and +9.23 points over the best baselines. | B. Supports structured planning and execution validation, especially for multi-table work. The composer carries most of the ablation gain. Removing the processor slightly improved the GPT-4o average while hurting weaker/multi-table settings. Five paper authors performed human annotation; the authors acknowledge evaluator bias, temporal errors, and incomplete semantic metrics. |
| DashChat | Industrial-dashboard pattern retrieval, a DSL, intent-specific parallel agents, explicit evaluation/repair, chat plus structured edit bubbles, and visual history/reversion. | Fifty held-out prompts: 100% executable, 94% exact spec consistency, and 41.4 s end-to-end; the two single-pass baselines reached 76–80% consistency and took 62.9–134.7 s. User study: 17 domain professionals and 11 designers; designers compared it with Tableau after a 30-minute tutorial. | A/B for rapid prototypes. Strong evidence for a constrained prototyping environment and iterative negotiation, not production analytics. It generates mock data; participants identified domain mismatch, acronym, direct-editing, and style-control limits. The Tableau comparison favors a prompt-first prototype task and says little about governed, production data work. |
| Lumen | Coordinator routes to SQL, Vega-Lite, Deck.gl, chat, source, table/document-list, and validation agents over serializable declarative pipelines and views. Deterministic profiling and cleaning precede some agent work. | Current, inspectable implementation and tests; no persuasive published comparative efficacy study found. A notable validation path fails open by assuming completion when structured validation cannot be parsed. | C. Strong mechanism source, not outcome evidence. Borrow serializable pipelines, real access-based roles, and profiling-before-aggregation. Reject fail-open semantic completion. |
| LIDA | Data summarization, persona-conditioned goal generation, chart code generation, six-dimension LLM evaluation, repair, and recommendation. | Influential open-source baseline; inspected revision has not moved since 2024-08. The evaluator is code/text-prompt based and repair echoes feedback into another generation call; it does not inspect the final rendered image or produce regression evidence. | C/D and historically useful. Keep as the canonical prompt-pipeline baseline. Reject its evaluator/persona pattern as current acceptance evidence. |
| PlotGen | Query planner and code generator followed by numeric, lexical, and visual feedback agents; the numeric agent de-renders the result with a VLM. | MatPlotBench 100: 65.67 with GPT-4 versus 61.16 for MatPlotAgent and 48.86 direct. Five participants reviewed 200 sampled requests; only 40.5% were completely accurate and 24.5% somewhat accurate. | B-, directionally useful. Supports multimodal feedback, especially lexical/visual checks, but lacks an equal-budget baseline, relies heavily on GPT-4V throughout, and contains reporting inconsistencies. It does not establish that multiple agents are the cause. |
| NL4DV-LLM | The model emits an inspectable analytical specification (attributeMap, taskMap, visList) and can return multiple interpretations for ambiguous prompts. |
740 queries over three datasets: GPT-4 prompt approach 87.02% versus rule-based NL4DV 64.05%, at roughly 25 s versus 3 s. Two authors graded, with a third tie-breaker; any valid ambiguous interpretation could count. | B, older-model evidence. Borrow explicit task vocabulary, mapping visibility, and multiple interpretations. Do not trust its generated confidence scores or equate valid syntax with correct attribute/encoding binding. |
| Data Formulator 2 | Concept binding: direct manipulation specifies encodings; concise natural language requests missing transformations. Data threads preserve branch/backtrack context. | CHI 2025 study described above. Participants used charts, transformed tables, code, and explanations differently to verify outputs. | A within a small reproduction study. This is better evidence for mixed initiative than for agent autonomy. |
Skill and plugin packages
| Package | Package class | Verification actually present | Outcome evidence | Assessment |
|---|---|---|---|---|
| Vizro end-to-end flow | Six skills split design, chart/layout selection, build, YAML, and actions; five required spec/test artifacts | AST checks for raw/unaggregated charts and color policy; required terminal inspection; Playwright walk of every page and every action; console, network 500, server traceback, screenshot/spec comparison, and test report | Three dashboard-build and four interaction eval prompts with explicit expectations; README says tested with two Claude 4.6 models, but no aggregate held-out result or independent user outcome found | C, strongest inspected skill mechanics. Borrow staged specs, targeted deterministic checks, action enumeration, browser evidence, and test receipts. Do not infer general chart quality from seven fixtures. |
OpenAI visualize-data |
Large workflow skill: question/takeaway first, chart contract, data sufficiency thresholds, surface routing, denominator/uncertainty/source rules, final-context rendering and inspection | Repository tests cover renderer, transform, tooltip, axis-domain, HTML fallback, and delivery contracts; the skill requires QA in the delivered surface | No held-out behavioral comparison of an agent with/without this skill found | C. Strong contract and coverage checklist. Borrow the chart contract and final-context QA. Treat prose thresholds as revisable defaults, not universal laws. |
| AntV chart visualization skills | Thin chart-image API skill plus deep G2/G6/X6 skills with strict version constraints and hybrid retrieval over a large reference corpus | Eval code contains structural/API checks, code similarity, a Playwright render tester, blank detection, and a VLM visual scorer | Included saved runs cover 174 G2, 97 G6, and 136 X6 cases with high structural similarity/hit rates. The July retrieval result files do not contain render results; “success” largely means generation completed without recorded structural failure. No no-skill baseline is included in those files. | C. Excellent evidence that version-specific constraints and progressive retrieval target real library hallucinations; insufficient evidence of visual correctness. Run a local with/without retrieval-and-render ablation before transfer. |
| SciVisAgentSkills | Version-pinned operational guides for napari, ParaView, Topology ToolKit, and VMD/MDAnalysis, including headless execution and render–inspect–adjust loops | Deterministic image, code, and rule checks plus multimodal judging across 108 expert-designed tasks | Claude Code and Codex were each tested three times with and without the relevant skills. Quality improved in all ten suite-by-agent comparisons, although one completion measure fell. | B, direct but bounded. This is the clearest visualization-specific skill ablation found. The authors built and evaluated their own packages; the tasks are scientific, and no independent reproduction or reader outcome was found. |
| Markdown Viewer Vega skill | Compact renderer adapter with syntax notes and examples | No task fixtures, data checks, or render loop found in the skill | None found | C/D. Useful as surface syntax, not a visualization method. |
data-visualization in claude-skillz |
Long single-file primer covering chart selection, Cleveland–McGill ordering, accessibility, performance, libraries, and layout | No scripts, fixtures, or evaluation harness found | None found | D. Good checklist specimen and prompt-only baseline. It packages advice but cannot show that an agent followed it or that a reader benefited. |
These packages make a useful maturity ladder:
static advice
-> versioned constraints and on-demand retrieval
-> explicit design/build artifacts
-> deterministic source/code checks
-> final renderer/browser inspection
-> interaction coverage and evidence receipts
-> held-out behavioral comparison with and without the package
-> human outcome study
SciVisAgentSkills reaches the paired-ablation rung for specialized scientific work, but not independent reproduction or a human outcome. Vizro reaches furthest on execution evidence; AntV has the largest included retrieval/code benchmark; the OpenAI skill has the broadest chart-contract and delivery QA. Those are different strengths and should not be collapsed into an install count or one “best skill.” The companion skill-package deep dive compares 18 files or families and separates registry installations from repository popularity and behavioral evidence.
Governed commercial environments
| Product surface | Current first-party mechanism | What it suggests | Evidence limit |
|---|---|---|---|
| Power BI Copilot | Builds a report page by selecting tables, fields, measures, and charts from a semantic model; generated pages remain editable with normal tools; answers can reference source visuals. | Bind generation to governed measures and retain direct author control. | Current capability documentation, not a comparative accuracy or user-outcome study. |
| Tableau Agent | Works within a connected data source and current worksheet state; uses field metadata and sampled values; creates/changes visualizations, calculations, filters, and sorts; dashboard Q&A entered beta in July 2026. | Keep agent scope close to existing authoring state and make direct manipulation the recovery path. | Tableau explicitly says to review results and treats the output as a starting point. It currently cannot choose a source, model data, build full dashboards in viz authoring, or create many interactions. |
| Looker Conversational Analytics | Grounds queries in LookML, permissions, descriptions, samples/fuzzy value search, custom instructions, business glossaries, and optional verified queries. The agent selects fields/filters while Looker composes database queries; optional Python handles advanced analysis. | A maintained semantic layer and verified examples are more reliable grounding than a generic analyst persona. Different domain agents can be policy/configuration packages over shared governed data. | Product documentation says outputs can be plausibly wrong and must be validated. No public evidence read here isolates which grounding feature improves end-user decisions. |
The commercial systems are especially environment-dependent. Their main advantage is not a universally better model. It is custody of semantic models, permissions, field metadata, verified queries, authoring state, and the native renderer. That advantage does not transfer to an open-file editorial workflow unless another system is given equivalent data contracts.
The evaluation frontier
DashArena materially changes the map
DashArena, published 2026-08-11, contains 234 open-ended tasks derived from high-quality Tableau Public dashboards across 14 clusters. A candidate returns a single-file ECharts dashboard and a structured two-turn interaction trajectory. A Playwright executor replays the trajectory and gives the pairwise judge task context, screenshots, schema, and execution evidence.
The strongest results are about the evaluation method:
- Six dashboard-experienced human annotators labeled 100 pairs, three labels per pair; 99 were evaluable. Human inter-rater agreement was only moderate (Fleiss kappa .384), and 52 pairs were unanimous. Dashboard quality is not a naturally objective scalar.
- The distilled DashJudge-8B agreed with humans 79.8% of the time (kappa .600). Removing interaction evidence reduced agreement to 71.7% (kappa .441), while deterministic rules alone achieved 42.4% (kappa .095). Intended-use evidence added 8.1 percentage points.
- A 100-trajectory audit found 99.8% valid targets and 96% all-pages coverage, but per-page control coverage was 80.8%. Of 841 actions, 97.3% were attributable and 93.3% matched the intended behavior; the other 6.7% exposed errors rather than hiding them.
- On 120 held-out tasks, the top model, GPT-5.5, scored 1449 versus a human baseline at 1321, with overlapping uncertainty. The authors explicitly treat the human dashboard as a practical anchor, not an oracle, because it predates the derived task.
- No model exceeded 86% render success or 74% trajectory replay success. In a targeted 100-failure audit, construction/runtime caused 26.9%, interaction 25.4%, data computation/binding 22.4%, presentation/readability 16.4%, and analytical issues 9%. Of 30 execution-clean failures, 21 still had semantic defects.
This qualifies the earlier conclusion that communicative-value evaluation was empty. DashArena now provides a serious, human-calibrated measure of task-grounded analytical support and interaction quality. It does not reverse the conclusion about readers. The benchmark does not observe whether a target audience comprehends an argument, learns a concept, remembers the message, or makes a better real decision. Its pairwise aggregate is also not a universal taste or audience model. The right update is “a major middle layer now exists,” not “communicative value is solved.”
The naming has outrun the instrument
Two further 2026 benchmarks reach toward communicative value and then measure something else. Reading them together is more informative than reading either alone, because they fail in the same direction from opposite starting points.
SciVisAgentBench is the most systematic scientific-visualization agent benchmark released so far: 108 expert-crafted cases over a four-dimensional taxonomy, a multimodal evaluation pipeline, and an IRB-approved validity study with 12 domain experts. Its taxonomy makes Scientific Insight Derivation a first-class visualization operation. But that is a task the agent performs, not a property of the artifact that is scored. Scoring is rubric agreement with a single expert-authored ground-truth visualization, judged 0–10 by an MLLM, plus PSNR, SSIM, and LPIPS similarity to that reference, plus token and time efficiency. Tasks are deliberately curated to admit one explicit outcome. The 12 experts are not readers being measured; they are raters whose agreement with the LLM judge is being measured. The paper is candid that open-ended goals with multiple valid outcomes “remain difficult to evaluate reproducibly at a benchmark scale,” and names the same obstacle from the inside: different visualizations may convey the same insight, which complicates automated scoring.
MultiVis-Agent goes further in language and no further in instrument. Its evaluation has an explicit high-level perceptual layer of six weighted dimensions — chart-type appropriateness, spatial layout, textual elements, data representation, visual styling, global clarity — described as “capturing human-perceived effectiveness.” The scorer is a vision-language model applying rubrics to a rendered image. No human reader appears anywhere in the evaluation.
The gap, stated precisely, is no longer that the field has failed to name communicative value. Both of these name it, or something adjacent to it. The gap is that naming has outrun measurement, and a model judge scoring a rubric is increasingly being asked to stand in for a reader. That substitution is convenient, reproducible, and cheap, and it is not the same evidence. A benchmark can report a high perceptual score for a chart that no human has read.
For anyone selecting a benchmark: read what the metric is computed on, not what the metric is called. A dimension named for clarity, effectiveness, or insight may be a model’s rubric score against a reference image. That is a useful signal about conformance and a weak one about communication.
Dashboard2Code exposes state that a screenshot can hide
Dashboard2Code reconstructs interactive Plotly Dash applications from screenshots, optional DOM, and interaction. Its benchmark combines 58 real-world seed dashboards with 122 generated examples, then checks visual fidelity, code, dynamic browser behavior, and 450 interaction tasks. Ninety generated dashboards were also scored by three visualization- experienced graduate evaluators; the final automatic metric correlated .781 with their ratings.
The most important result is not one total score. DOM access sharply improves exploration, complex callbacks remain harder, and systems can produce a visually responsive but factually wrong state. A static image judge cannot see that failure. The scope boundary matters just as much: Plotly Dash only, fixed desktop viewport, no animation or popups, and no reader study.
Chartography makes professional reading a separate capability
Chartography, released 2026-08-11, contains 100 difficult professional chart-reading tasks across 12 domain labels. Each task was authored by a practitioner and independently verified by three experts, including an acceptable answer range. The best of 30 frontier configurations reached 45.0% mean pass@1 over twenty trials per task.
Higher reasoning effort usually helped, but not enough to erase the main failure: sparse axes, 3D projections, contours, and domain conventions can be misread before reasoning begins. More tokens then produce a longer explanation of the wrong visual premise. The benchmark is adversarially difficulty- screened, so it does not say models fail on 55% of ordinary charts. It does say that a general “chart literacy” score is too broad a release gate for professional, consequential reading.
Raiven narrows generation failure in a particular environment
Raiven’s .988 VMPC is strong counterevidence to treating direct-code chart generation failure rates as immutable. A formal DSL plus compiler can nearly eliminate many specified-mark, encoding, linking, hallucination, and execution failures. But the benchmark fully specifies the target visualization, and the largest delta occurs in 3D/scientific tasks where generic web code generation is weak. In ordinary 2D information visualization, direct frontier models nearly match Raiven on its own metric.
The transferable claim is conditional: when an environment has a stable visual grammar, recurring hard syntax, and deterministic rendering, invest in the representation/compiler. It is not evidence that a DSL can choose the right question, identify a misleading comparison, or replace final human judgment.
Benchmark construction is part of the technique
The companion benchmark crosswalk documents the input object, requested output, source distribution, grader, human context, supported claim, and nonclaim for the benchmarks carrying the main conclusions in this review. It reveals a measurement progression:
- exact or near-exact code/spec matching;
- syntactic legality and render success;
- data/encoding correctness against a reference;
- chart extraction, visual grounding, question answering, and comparison;
- integrity or quality judgment with human calibration;
- interaction state and task-grounded replay;
- actual user performance, comprehension, or decision outcome.
Many impressive percentages in tool repositories live at levels 1–2. nvAgent and Raiven reach level 3 in bounded ways. ChartQAPro, Chartography, ChartDiff, and the two 2026 PolyChartQA benchmarks expand level 4 across realistic, professional, comparative, multi-chart, and multilingual cases. Misviz, VisJudge, and misleading-chart robustness studies measure parts of level 5. Dashboard2Code and DashArena reach level 6. Data Formulator 2 and DashChat provide small, contextual evidence at level 7 for authoring/prototyping experience, not for reader comprehension or decision quality. No one result spans the stack.
Three construction effects now deserve the same attention as model scores. Human-authored multi-chart questions were up to 27.4 percentage points harder than model-generated questions; ChartDiff’s lexical-overlap metrics favored systems that its human-aligned judge rated much worse; and misleading-chart system rankings changed between controlled synthetic and heterogeneous real charts. A benchmark is a task, distribution, and grader—not a neutral name.
Five validity threats block a configuration verdict
The latest audit asks a narrower question than the landscape: does this evidence authorize one Vizier configuration? The answer is no. It records five distinct boundaries rather than hiding them in one confidence label:
| Threat | Current result | What must happen next |
|---|---|---|
| Cross-literature join | retained unresolved | test the visualization-specific architecture under equivalent current receipts |
| Flattering-prior risk | partly tested, retained | preserve the hostile pass’s failed hypotheses and preregister rivals before inference |
| Benchmark-to-quality inference | rejected | require a new receipt at every later evidence layer |
| Recency boundary | bounded | keep durable task constructs separate from perishable model/harness results |
| Single-configuration limitation | retained unresolved | run the comparison; prepared packets and verifiers are not results |
Two direct human/model comparisons make the benchmark boundary concrete. CHART-6 evaluates eight models on 851 items from six visualization-literacy assessments, ten times each. The tested models underperformed the held human data on average, and every item-level model error pattern remained below the human noise ceiling. A newer high-level interpretation study compares 24 people’s descriptions with three multimodal models over 60 charts. Models matched the study’s bounded analytical designer intent more often, yet tended to enumerate structure and values where people more often built narratives and used visual scaffolding.
Those findings are not inconsistent once task, model, prompt, scoring, and vintage remain attached. They show why aggregate task success, human-like errors, analytical-intent matching, and reader experience are separate claims. The public rule is therefore strict: a benchmark licenses only its named task-and-grader lane. Overall quality, audience fit, comprehension, accessibility, accepted delivery, maintenance, and a configuration winner each require new evidence.
Research gap ledger
An open question is not the same as an empty field. Some gaps now contain a bounded study, some contain evidence that a failure occurs, and some still lack the measurement needed to make a decision. The status language below is literal:
- Measured but bounded means a public comparison exists but does not yet generalize to ordinary production.
- Failure demonstrated means research establishes that the failure occurs, not its population frequency or best remedy.
- Partly answered means a human study covers one task, audience, or time window while leaving consequential transfer questions open.
- Open means the reviewed corpus does not contain direct outcome evidence for the question as posed.
| Gap | What the evidence establishes now | What remains missing | Next evidence that would change the status | Status |
|---|---|---|---|---|
| One-pass competence on realistic work | RealChart2Code and DashArena show that frontier systems can produce materially faithful charts and open-ended interactive dashboards, while still missing render, replay, and semantic requirements. | The share of ordinary editorial, operational, scientific, and governed-BI work that is acceptable without repair. | A stratified held-out task set sampled from real work, scored through data, render, interaction, and human acceptance. | Measured but bounded [Q9] |
| Capability-to-practice bridge | Ten primary cases were audited for technical-to-human joins. HAIChart links recommendation performance to 17-person controlled use, interactive task decomposition links system error to correction across 108 analyst episodes, and ChartAttack carries generated attacks into a 48-person reader experiment. | An immutable system/version plus frozen acceptance contract and real delivery. No audited case continues through a complete reader or decision outcome, later maintenance, and whole cost. | Carry one controlled result forward without changing artifacts, tasks, configuration, or version: accept under a frozen rule, deliver it, observe representative use, and return after a material event with full cost. | 3 controlled partial · 0 accepted-delivery |
| Data and semantic fidelity | nvAgent, Raiven, Text2Vis, and DashArena test bindings, calculations, or task-grounded semantics in bounded environments. DashArena found semantic defects even among execution-clean outputs. | Accuracy against organization-owned measures, ambiguous fields, changing sources, permissions, and unstated local rules. | Public field evaluations stratified by semantic-model quality, task ambiguity, and data conditions. | Measured but bounded |
| Correction without regression | RealChart2Code demonstrates regressive editing: a requested fix can introduce new faults in previously correct code. The novice study found 11 of 16 clutter fixes and 8 of 9 unusable-chart repairs failed. | Accepted-artifact correction cost, abandonment, regression after delivery, and reliable stopping rules across models and environments. | Multi-turn studies that retain every state, classify introduced and repaired faults, and follow work through acceptance. | Failure demonstrated [Q15] |
| Critic architecture and specialized vision | VisJudge-7B outperformed the tested general models on its expert-adjudicated quality rubric; ChartAgent’s chart-specific tools materially improved the same general reasoner on numeric QA; chart-aware grounding and OCR models add useful perception. New transfer tests also show strong general models overtaking older chart specialists. | A routing rule, same-generator and equal-budget end-to-end critique lift, cost, interactive and mobile coverage, source-fidelity checks, independent reproduction, and reader outcomes. | Evaluate deterministic checks, current general vision, narrow specialists, and composed critics on the same generated artifacts and consequential defect set. | Partly answered [Q7] [Q16] |
| Readability versus integrity | Frontier-model visualization literacy and misleading-chart detection can rank differently, showing that decoding and integrity are distinct capabilities. | Whether that divergence reproduces on current generated artifacts and predicts actual human misreadings. | A shared chart set scored independently for data truth, deceptive encoding, readability, and reader outcomes. | Open [Q8] |
| Learning, authoring outcomes, and expertise | The human-skills synthesis anchors expertise in consumption, construction, critique, and connection plus data, domain, tool, situated-judgment, and delivery resources. One randomized study supports immediate post-removal comprehension from proactive scaffolding; adjacent trials show assisted performance can outrun learning; a no-AI visualization study measures ordinary decay over 12 months. A protocol now crosses conventional, answer-oriented, and metacognitive practice with immediate withdrawal, six-week and six-month unfamiliar transfer, and a separate blinded reader stage. | A stable definition of “novice”; approved and powered crossed capability/accessibility cells; a frozen model and harness; delayed visualization construction and far transfer; real-reader delivery on own devices and assistive paths; and effects over weeks or months. No captured longitudinal study causally estimates AI-driven visualization learning or atrophy. | Approve, register, and run the protocol while keeping access, assisted performance, learning, cost, accessibility, and reader outcomes separate. A protocol verifier is not a human result. | Partly answered; study designed, not run [Q13] |
| Reader comprehension and decisions | A ten-row AI-role audit finds one direct controlled effect where selected AI-generated ChartAttack outputs reach independent readers. Lexara adds a distinct longitudinal middle: six CVA developers used a deployed evaluation toolkit for two weeks across 38 experiments, 57 newly authored cases, ten models, and six prompts. Three production afterlives and the later same-feature InfoViP field-use and operating line bring the audit to 15 named cases. | Zero case binds an immutable artifact/build and exact configuration to accepted intended-reader delivery, consequential decision or calibrated trust, whole cost, and later post-release same-lineage recheck. Development selection, a bounded voluntary field evaluation, and at-scale processing do not fill those states. | Keep twelve receipts on one lineage from AI role through artifact/build/configuration, acceptance, delivery, context, decision, trust, elapsed event, recheck, maintenance authority, and cost. Carry one evaluation or operational record into a frozen release and affected-audience return rather than borrowing receipts across cases. | 15 named · 1 longitudinal evaluation · 1 counted field-use + operating near-miss · 0 full episodes [Q10] [Q14] |
| Decision-visualization review lineage | A 2025 PRISMA review searched five databases through 1 July 2024 and retained 127 studies: 118 empirical papers and nine reviews. Its supplement prints 126 citations plus ?; the source-internal B59 overlay repairs that placeholder. A complete six-domain crosswalk now binds 122 of the 127 supplement positions to distinct A208 review keys. Of those keys, 114 carry DOIs, two carry confirmed PubMed identifiers, five expose no publisher identifier, and one carries a printed PMID that resolves to another paper. |
Five printed rows—Jiao (2022), Burt et al. (2017), Theis et al. (2018), M. Lu (2020), and Islam et al. (2022)—have exact-looking external DOI candidates but no review-controlled join, so none enters the authoritative frame. The conflicted B13 PMID also remains quarantined. Every lifecycle field is still unaudited; the qualifying count is unknown, not zero of 127. | Preserve source label, authoritative review key, external candidate, identifier validation, and admission authority separately. Obtain review-author, publisher, extraction, or A208-controlled joins for the five gaps, repair or remove the conflicted PMID, and only then audit the stable frame against the eight receipts—or capture one direct same-artifact GenAI episode independently. | 127 positions · 122 review keys · 5 authority gaps · denominator unknown |
| Artifact custody and audience delivery | A bounded title-cue screen over the 122 authoritative A208 keys surfaces one explicit GPT study. Beyond Generating Code evaluates 91 quiz questions and nine homework assignments; its one-commit supplement pins 1,271 blobs spanning prompts, outputs, generated code, and grader bundles. | The runner exposes model aliases rather than an immutable provider snapshot and complete configuration. The paper does not directly evaluate the final project. Government, UN, and primary-school audiences appear in prompts, not as recipients or evaluators. No named release, later commit, audience delivery, outcome, or recheck is exposed. | Carry the pinned bundle into a frozen release and acceptance rule, deliver the accepted artifact to its named audience, retain outcomes and whole cost, and return after a material event. Treat the other 121 title nonmatches as not surfaced, not excluded; the five unresolved review positions and full lifecycle denominator remain null. | 122 title-triaged · 1 GPT signal · 1 pinned-output bundle · 0 delivery/recheck chains |
| Mature delivery and afterlife without an AI role | Exact-DOI discovery returns all 114 authoritative DOI records; 87 expose abstracts. AIDSVu is the sole abstract with explicit visualization, duration, observed-use, decision, and audience cues. Its paper reports ten years of public delivery, 501,527 unique users in 2019, named planning applications, and recurring governance. A 2026 data release and current FAQ add a later event and annual update/limitation contract. | The held surfaces expose no AI visualization role, immutable artifact/build, version-bound measured audience or decision outcome, calibrated trust, affected-audience recheck, or whole cost. Aggregate reach and self-reported use examples cannot be borrowed into an AI feature. The 86 GenAI cue nonmatches are not exclusions; 27 DOI rows lack abstracts and eight stable keys lack DOIs. | Keep AIDSVu as a non-AI comparator for durable public-data custody, audience surfaces, governance, and afterlife. For an AI claim, bind one immutable model/configuration and accepted artifact to intended-audience delivery, measured consequence, later material change, and affected-audience return. | 114 DOI records · 87 abstracts · 1 non-AI afterlife · 0 complete AI lifecycles |
| Primary human evidence below the abstract layer | All 27 no-abstract DOI rows have now been investigated. Six have full primary texts, 20 have primary abstract or official content, and one remains content-unassessed. The held cases separate controlled audience measurement, public-sector co-design/demo/bounded use, enterprise production intent, expert redesign, practice context, risk-ranking decision support, and a public non-AI delivery afterlife. | The explicit GenAI and AI-visualization-role screens are empty for the 26 substantive rows. No row joins an accepted field release to a consequential affected-audience outcome and later recheck. B91 remains access-blocked and unassessed; task evidence, method description, co-design, demo exposure, production intent, expert review, practice application, and public reachability belong to different artifacts and stages. | Pause generic B91 retrieval after three bounded passes and reopen only on an exact publisher abstract/body, accepted manuscript, correction, or new lawful deposit. Otherwise pursue the full B92 paper or one exact AI-assisted delivery-to-recheck episode. | 27 investigated · 6 full · 20 abstract/official · 1 access-blocked/unassessed · 0 generic B91 targets |
| Authority repair below the DOI layer | The eight stable non-DOI keys now yield three exact DOI repairs, two publication-year corrections, and one rejected foreign PMID. Custody separates four full texts, two primary abstracts, one issue excerpt, and one gated primary. The EMR cancer diary supplies the strongest practice receipt: increasing system-log use plus an independently recovered 11-clinician QUIS result with median satisfaction 4.38. | Seven held content screens expose no explicit GenAI cue; gated B81 remains unassessed. The cancer diary exposes no AI role, immutable accepted build, log denominator, patient or decision outcome, later event, or affected-clinician recheck. The complete-review denominator remains null. | Apply repaired identifiers without rewriting A208. Preserve theory, prototype, simulation, technical experiment, embedded use, accepted release, outcome, and recheck as separate stages. Recover B81, B112’s full paper/log denominator, partial B19/B142 text, or one of the 14 content-unassessed DOI rows. | 8 keys · 3 DOI repairs · 1 operational comparator · 0 AI lifecycles |
| Human evaluation in the remaining DOI layer | Nine high-signal rows were investigated; eight now expose a primary abstract or official project record and seven explicitly report participants or evaluators. InfoViP is the delivery near-miss: seven FDA safety evaluators supplied requirements, evaluated the prototype, and had suggestions addressed. An official FDA project page describes NLP and unsupervised learning and says an enhanced version will be installed in production. | “Will be installed” does not confirm installation, acceptance, routine use, regulatory outcome, later change, or evaluator return. The eight content surfaces are not full papers. A subsequent residual pass reduces the content-unassessed DOI remainder from 14 to three. None of the eight contains an explicit GenAI cue. | Seek an installation, acceptance, use, regulatory-impact, later-event, or evaluator-return receipt for InfoViP; recover full methods for the evaluated rows; and preserve evaluated, planned, installed, accepted, used, outcome-bearing, and rechecked as separate states. | 9 investigated · 8 content surfaces · 7 human evaluations · 0 lifecycles |
| AI production component, QA afterlife, and operating-authority boundary | A six-month internal real-work InfoViP evaluation records 20 unique reviewers and 58 submissions. The final CIOMS report says the component was approved for historical/live ETL, installed in AWS, integrated with AERS, and by 2025-07-30 had screened >30 million historical plus ~8,000 daily submissions. Case-series use is fully implemented and administrators monitor the pipeline/output. The now-closed Elsa/API/UI opportunity names Joshua Xu, Leihong Wu, and Oanh Dang as research mentors. | Closure supplies no selection, onboarding, work, acceptance, or release receipt. The opportunity says a fellow is a nonemployee barred from inherently governmental functions. Official profiles support bounded research and project-lead roles, not InfoViP service operation, maintenance, QA execution, release approval, or authorization. An adjacent AI pharmacovigilance QA project describes a literature review and planned report but does not name InfoViP or a completed audit. The current HHS inventory still records implementation N/A and no ATO, and the Elsa 4.0/HALO launch never names InfoViP. |
Route research questions to the three mentors, but keep operator, maintainer, QA executor, release approver, authorization authority, immutable build, completed audit, application join, users/outcomes, return, and whole cost as separate required receipts. Treat opportunity closure as an administrative state, not execution. | 3 research mentors · 0 operating authorities · opportunity closed · no InfoViP-Elsa join · 15 named · 0 full |
| Residual DOI primary recovery | All 14 residual rows were investigated; eleven gained substantive primary content, including one full chapter that connects three UX experts and 25 unique problems to an implemented third version. A corridor stakeholder case and university-network case add practice context. | At that pass, B92, B91, and B110 were metadata-only. No row exposed an explicit GenAI visualization role, accepted field release, routine-use denominator, consequential affected-audience outcome, later material event, or affected-actor recheck. Display context, expert review, stakeholder application, and multi-unit evaluation are not deployment. | Preserve this historical 14 → 11 + 3 receipt, then apply the later B110 recovery without rewriting the earlier batch. Keep every receipt on its own artifact lineage. |
14 investigated · 11 substantive · 1 full text · 3 unassessed at that pass |
| Final-three content and delivery recheck | The exact three rows were rechecked. B110 now has an official bibliographic abstract, an immutable framework state, a pinned DiscoverWater application, and a same-named KU interface that is reachable in August 2026. | At that pass B92 and B91 remained content-unassessed. B110’s live bytes are not bound to a SHA, and no receipt shows acceptance, continuous or ordinary use, consequential stakeholder outcome, accessibility acceptance, affected-audience return, or whole cost. Its held surfaces expose no AI visualization role. | Treat source custody, a current public route, accepted build, ordinary use, audience outcome, and recheck as separate states. Bind live bytes to a release and measure named-audience use before making an adoption or impact claim. | 3 rechecked · 1 newly substantive · 1 live surface · 2 unassessed at that pass |
| Live-build lineage and audience receipt audit | A deterministic reconstruction proves that the dated live DiscoverWater page does not match the sole published v1.2 page state; the linked v2.0 source is R/Shiny. Partial data lineage remains: eight of twelve live dependency basenames occur in the pinned application tree, and two of three sampled asset pairs are canonically equal. An official 2018 record adds a limited prototype demonstration. | No public manifest binds the live page to an immutable build. Analytics and comment hooks supply instrumentation, not traffic or audience evidence. Acceptance, ordinary use, consequential outcome, accessibility acceptance, later affected-audience return, whole cost, and an AI role remain missing. | Preserve the dated page digest, source reconstruction, dependency coverage, and audience receipts separately. Reopen on an exact deployment manifest or version-bound audience record—not another same-name source link. | 1 live page · 0 exact builds · 8/12 names · 2/3 canonical samples · 0 audience outcomes |
| Final-pair lawful primary recovery | B92 now has an exact publisher abstract describing a multicriteria hydrogen-pipeline risk model, Monte Carlo simulation, Kendall’s tau rank comparison, and graphs for ranking sections and targeting mitigation. Three bounded B91 passes cover six named lawful surfaces. | The abstract names no human sample or AI visualization role and supplies no release, intended-audience delivery, observed use, consequential outcome, later recheck, or whole cost. The full B92 paper is unheld. B91 remains content-unassessed; paused access is not a negative or world-level absence. | Reuse B92 only as a non-AI uncertainty-and-ranking method comparator. Reopen B91 only on exact new lawful custody; otherwise move effort to the full B92 paper or one exact AI-assisted delivery-to-recheck episode with whole cost. | 1 publisher abstract · 1 access-blocked/unassessed · generic search paused · 0 AI lifecycles |
| International and non-English use | Multilingual POLYCHARTQA now tests 22,606 charts and 26,151 QA pairs across ten languages and finds substantial English/non-English gaps, especially in visual-language transcription. MM-JudgeBench finds language-dependent accuracy and bias across 25 languages, including a chart-centric judge subset. | Real non-English authoring and reading, culturally situated chart conventions, code-switching, locally used analytics environments, more low-resource languages, and reader outcomes. Both benchmarks translate English-centric source material rather than sampling ordinary local work. | Human-authored multilingual tasks and delivered-reader studies sampled across languages, chart conventions, institutions, and devices, with source distributions reported separately. | Partly answered |
| Skill-package effectiveness | Current packages provide advice, versioned references, contracts, executable checks, or browser inspection. SciVisAgentSkills improved quality in all ten paired suite-by-agent comparisons across 108 scientific-visualization tasks, although one completion measure fell. | Independent tests of popular generic, dashboard, accessibility, mobile, and explanatory-visualization packages; decay after model releases, negative transfer, reader outcomes, and total instruction cost. | No-skill, prose-only, narrow specialist, retrieval, verifier, and browser-evidence ablations repeated across current models, harnesses, and environments. | Partly answered [Q12] [Q17] |
| Context-sensitive routing | Established task and environment typologies explain why newsroom explanation, governed BI, open exploration, science, education, and operations impose different obligations. | Which context dimensions actually change the best generator, critic, evidence bundle, or human gate—and which can safely remain shared. | A factorial evaluation that varies audience, purpose, stakes, data custody, interaction, and delivery while holding the task family stable. | Open [Q11] |
| Interaction and delivered state | DashArena replays model-authored trajectories and materially improves human agreement by showing interaction evidence. Dashboard2Code tests callbacks and hidden state across 180 fixed-desktop Plotly Dash applications and 450 interaction tasks. | Real-user exploration, responsive and mobile layouts, animation, popups, keyboard paths, authenticated applications, permissions, exports, assistive technology, device variation, and post-deployment state. | Multi-viewport browser and human studies against the actual delivery surface, including state transitions, failure recovery, and non-happy paths. | Measured but bounded [Q15] |
| Maintenance and total cost | Twelve cost fragments are held. ChartAgent adds a quality/tool-call proxy; DV-World adds interaction charge; Selective TTS matches a declared partial inference budget. VisCoder2 adds a same-endpoint retry topology: over 888 tasks, VisCoder2-32B and GPT-4.1 both finish at 732 execution passes after 584 versus 714 revisions. The released debug paths discard the comparable usage and latency telemetry needed to price that difference. Ledger v7 separates matched declared partial budget, conditional attempt topology, observed route-wide use, and accepted-artifact cost. | Zero of twelve held fragments has both equivalent observed per-arm route-wide cost and a frozen accepted-artifact denominator. Execution pass is not human or production acceptance, and no same-task record also joins preparation, delivery, a later event, independent recovery, maintenance authority, accessible-reader use, decisions, or calibrated trust under a current direct-work comparator. | Freeze one multidimensional cap and acceptance contract; retain cold start through synthesis, effective topology/runtime state, calls/retries, tokens/image units, evaluator/tool work, accelerator use, latency, charges, human minutes, and accepted, rejected, abandoned, quarantined, and no-output states for every arm. Report the retry survival process and declared budget, observed use, eligible, attempted, candidate, accepted, delivered, and reader-successful ratios separately. | Partly answered; 0/12 accepted-cost comparisons [Q15] |
| Same-artifact lifecycle joins | Three dashboard cases—Prism, OpenClaw, and KubeStellar—cross delivery into later events. MAIDR adds a study surface, version floor, and later changes; Graphy adds repeated co-design. A six-receipt audit tests those five lines with PM4Py-UCM and DV-World. Every required receipt appears somewhere, but zero of seven rows joins repeated representative use, immutable tested build, exact exposure model, versioned release, later material event, and representative or actor-separated post-change recheck. | One exact lineage clearing all six receipts, then extending into independent recovery, maintenance authority, accessible use, decisions, calibrated trust, or whole cost. The current null is scoped to named paper, repository, release, issue, and follow-up surfaces through 15 August 2026. | Publish immutable exposure coordinates and re-run representative users after a named change on the released version. Preserve partial cells and search surfaces; never combine stages from different cases into a synthetic lifecycle. | 0/7 full release-validation joins |
| Production recovery states | The three dashboard afterlives all cross delivery into a later event. KubeStellar and Prism reach accepted repair, corrected delivery, and maintainer recheck. Zero case has an affected-user or independent-operator post-correction recheck, transferred maintenance authority, or all twelve states. | The first same-case affected-actor return receipt or observed exercise of second-person release or incident authority. OpenClaw issue 30/PR 31 and Prism issue 81 were rechecked through 16 August 2026 UTC. | Preserve actor, artifact/version, route, event, diagnosis, repair, delivery, maintainer recheck, affected-actor recheck, prevention, contribution, and authority as separate receipts. Issue closure is not recovery; contribution is not authority. | 3 afterlives · 2 restorations · 0 independent recovery · 0 authority transfer |
| Reproducibility and transfer | Several papers release code, data, or supplementary materials; others retain task sets, executors, model versions, or judges. Strong scores often depend on a renderer, grammar, or evaluator that does not transfer automatically. | Independent reruns, cross-renderer tests, stable public tasks, and calibration against readers or domain experts. | Full task, executor, judge, model-version, and failure-trace releases followed by third-party reproduction. | Open |
This ledger changes the research agenda in two ways. First, it prevents a single new paper from being narrated as “the gap is solved”: the relevant row moves only as far as the evidence permits. Second, it makes research pursuits testable. A skill-package survey must eventually lead to an ablation, not a popularity ranking. A specialized-vision survey must identify a same-task, equal-budget critic comparison, not merely another chart-question-answering leaderboard.
Active refresh triggers
The following releases or studies would materially change one or more rows:
- DashArena releasing its task set, executor, and judge for independent reproduction, or adding coding-agent, non-Tableau, numeric-fidelity, and real user-outcome tracks;
- Raiven releasing code and third parties reproducing its scientific- visualization results with independent graders and exploratory tasks;
- a public equal-budget visualization study isolating role/tool decomposition from extra inference;
- independent current-model evaluations of generic, dashboard, accessibility, mobile, and explanatory-visualization skills against the normal harness;
- same-task, equal-budget routing comparisons among deterministic checks, current general vision, and narrow chart specialists on end-to-end critique;
- a Graphy workshop-to-commit and model receipt, or a versioned Graphy release with formal or in-situ representative recheck after a material change;
- repair of the unidentified Education row in the 2025 decision-visualization review’s supplement, followed by a stable-ID row audit against GenAI role, immutable artifact, delivery, decision or calibrated-trust outcome, and later-recheck fields;
- reader studies measuring comprehension, retention, trust calibration, accessibility, mobile use, or consequential decisions from generated visualizations;
- current Data Formulator studies covering governed sources, persistent agent threads, and free exploration rather than reproduction;
- commercial BI vendors publishing reproducible error rates or user outcomes stratified by semantic-model quality and task type; and
- evidence that small, local deterministic contracts fail to transfer across the renderer and publication environments a system actually serves.
Sources inspected
Pinned repositories
- Microsoft Data Formulator at
5d4f7b3, MIT. - Microsoft LIDA at
d892e20, MIT. - HoloViz Lumen at
7d587f3, BSD-3-Clause. - McKinsey Vizro at
b1d7b11, Apache-2.0. - AntV chart visualization skills at
b47f2fe, MIT. - Markdown Viewer skills at
a3afd45; no repository license identified in the captured metadata. - OpenAI role-specific plugins at
fe5608d, MIT. - NTCoding
claude-skillzat21c6101; no repository license identified in the captured metadata.
Papers read in full for this update
- Wang and Deng, DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation, 2026.
- Irger et al., Raiven: LLM-Based Visualization Authoring via Domain-Specific Language Mediation, 2026.
- Shi et al., NL2Dashboard: A Lightweight and Controllable Framework for Generating Dashboards with LLMs, 2026.
- Ouyang et al., nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent Workflow, ACL 2025.
- Shen et al., DashChat: Interactive Authoring of Industrial Dashboard Design Prototypes through Conversation with LLM-Powered Agents, 2025.
- Wang et al., Data Formulator 2: Iterative Creation of Data Visualizations, with AI Transforming Data Along the Way, CHI 2025.
- Sah et al., Generating Analytic Specifications for Data Visualization from Natural Language Queries using Large Language Models, 2024.
- Goswami et al., PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback, 2025.
Human-outcome studies added by the gap-ledger pass
- Kuo et al., Vibe Visualizing: Exploring the Human-AI Interaction Dynamics in AI-Assisted Visualization, 2026 preprint: 20 novices, 60 sessions, and 175 charts.
- Kwon et al., Visualizationary: Longitudinal Critique of Visualization Design, 2024 preprint: 13 designers working on self-selected visualizations over three to five days, plus three expert raters.
- Yan et al., The Effects of Generative AI Agents and Scaffolding on Visual Analytics Comprehension, 2024 preprint: randomized 117-person comparison of data stories, passive GenAI, and proactive scaffolded GenAI.
- Palani and Setlur, Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual Analytics, CHI 2026: 22 developer interviews, 16 professional end-user observations, and a two-week field deployment with six developers, 38 experiments, 57 newly authored cases, ten models, and six system prompts.
- Ardito et al., AI-Supported End-User Development for Data Visualization, 2025: one-company exploratory study with eight interviews and three design-probe sessions.
Benchmark and training studies for the moving-baseline analysis
- Verma et al., CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models, 2025: eight models, 851 common items, six human-designed assessments, and item-level human error-pattern ceilings.
- Jeon et al., How Do LLMs See Charts? A Comparative Study on High-Level Visualization Comprehension in Humans and LLMs, 2026: 60 charts, 24 human participants, three multimodal models, and three prompt constraints.
- Wu et al., Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots, Findings of NAACL 2025.
- Rahman et al., Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text, EMNLP 2025.
- Zheng et al., CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution, ACL 2026.
- RRVF: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback, anonymous ICLR 2026 submission. Used only for its clearly labeled same-benchmark historical table and treated as preprint evidence.
- Li et al., Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards, 2026 preprint: 180 dashboards, 450 interaction tasks, and a 90-dashboard human validation of the evaluator.
- Chartography: A Benchmark for Professional Chart Understanding, 2026 preprint: 100 practitioner-authored, three-expert-verified tasks and 30 model configurations.
- Xu et al., POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering, ACL 2026: 22,606 charts and 26,151 QA pairs across ten languages.
- Ymyang et al., Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based Documents, ACL 2026: 4,738 expert-validated QA pairs across chart-bearing real-world documents.
- Ye, ChartDiff: A Large-Scale Benchmark for Comprehending Pairs of Charts, ALVR 2026: 8,541 chart pairs and a direct lexical-versus-human-aligned metric comparison.
- Efat et al., Beyond Single Plots: A Benchmark for Question Answering on Multi-Charts, Findings of ACL 2026: 534 multi-chart figures and separate human- and model-authored questions.
- Shu et al., FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models, ACL 2026: 1,200 real financial charts, 7,016 questions, and 26 evaluated models.
- Tonglet et al., Protecting multimodal large language models against misleading visualizations, ACL 2026: controlled vulnerability and six-intervention study over misleading and non-misleading charts.
- Laskar et al., Lost in Translation: Do LVLM Judges Generalize Across Languages?, Findings of ACL 2026: multilingual multimodal judge evaluation with a chart-centric subset.
Reader and accessibility study added by the open-question pass
- Touching or Chatting: LLMs and Tactile Charts for Blind and Low-Vision Chart Learning, 2026 preprint: 12 participants and 263 substantive queries; strong multimodal preference and spatial-model benefit without measured accuracy improvement.
Foundational task and environment frameworks
- Brehmer and Munzner, A Multi-Level Typology of Abstract Visualization Tasks, IEEE TVCG / InfoVis 2013.
- Segel and Heer, Narrative Visualization: Telling Stories with Data, IEEE TVCG / InfoVis 2010.
- Sarikaya et al., What Do We Talk About When We Talk About Dashboards?, IEEE TVCG / InfoVis 2018.
Current first-party product documentation
- Power BI: use Copilot with reports and semantic models, last updated 2026-05-26 when inspected.
- Tableau Agent FAQ, including July 2026 dashboard beta and current limitations.
- Looker Conversational Analytics overview, including LookML grounding, data agents, verified queries, and Advanced Analytics.
Model and harness milestone records
- GPT-4V system card, 2023-09-25.
- GPT-4o announcement, 2024-05-13.
- Claude 3.5 computer use beta, 2024-10-22.
- Claude Code research preview, 2025-02-24.
- Responses API and agent tools, 2025-03-11.
- GPT-5 for developers, 2025-08-07.
- Hosted computer environment and reusable agent skills, 2026.
Earlier evidence retained in the comparison
The underlying research packet also contains VisEval, MAST, Cleveland–McGill MLLM testing, Does It Run, VIS-Shepherd, Text2Vis, CoDA, RealChart2Code, equal-budget single-agent results, synthetic-persona studies, and LLM visualization literacy work. Those sources continue to control the claims about reader simulation, integrity versus readability, equal-budget coordination, and benchmark coverage. This update does not silently promote abstract-only captures from that packet to full-text evidence.
Update log
-
2026-08-18 — Named the gap more precisely: the naming has outrun the instrument. Two 2026 benchmarks that had been seen by title only were read in full. SciVisAgentBench makes “Scientific Insight Derivation” a task in its taxonomy but scores rubric agreement with one ground-truth image plus PSNR/SSIM/LPIPS, and its 12-expert study measures human–LLM judge agreement, not comprehension. MultiVis-Agent describes a perceptual layer “capturing human-perceived effectiveness” and instruments it with a VLM scoring six rubric dimensions, with no human reader in the evaluation. The standing conclusion is unchanged but sharper: the field now names communicative value and measures conformance, and a model judge is being asked to stand in for a reader. Evidence cut moved to 2026-08-18.
-
2026-08-16 — Retired a repeated retrieval route without promoting an access gap. Three bounded B91 passes now support a named-surface stop rule: the
6 / 20 / 1 / 0content ledger is unchanged, one row is access-blocked and unassessed, zero generic B91 targets remain active, and four exact conditions can reopen it. -
2026-08-16 — Mapped InfoViP research contacts without inventing operating authority. The closed Elsa/API/UI opportunity names three research mentors but supplies no execution receipt; its nonemployee boundary cannot furnish governmental approval. An adjacent AI-QA project does not name InfoViP or a completed audit. Operation, maintenance, QA execution, release approval, authorization, and the application join remain unassigned.
-
2026-08-16 — Repaired InfoViP’s chronology and application boundary. A future-tense FDA page is explicitly content-current to 2024, a 2026 FDA biography names a project lead without operations authority, and the Elsa 4.0/HALO launch contains no InfoViP join. Search recency and agency-level release no longer distort the maintained episode.
-
2026-08-16 — Final InfoViP report adds a production component and its QA afterlife. Added processing approval, AWS/AERS integration, >30 million historical plus ~8,000 daily submissions, bounded routine case-series use, and monitoring. Kept the absent QA plan, planned audits, incomplete roles, uninvestigated downstream effect, current no-ATO record, and all complete episodes visible. A 2026 Elsa/API/UI opportunity is prospective.
-
2026-08-16 — InfoViP gained counted field-use and award-envelope receipts. Added 20 unique reviewers and 58 submissions from a six-month internal real-work evaluation, a later 29-million-history plus daily operating pipeline, and three FDA awards totaling $2.40 million obligated / $2.86 million estimated. Kept acceptance/ATO, current routine use, outcome, affected-reviewer return, whole cost, ROI, and all complete episodes absent.
-
2026-08-16 — InfoViP added an operational-pipeline and governance- counterevidence row. Joined co-design, an operational pilot, 28-million- report plus incoming processing, and a later performance report while the current inventory still says development, implementation N/A, and no ATO. The exact-artifact audit is now 15 named cases and zero complete episodes.
-
2026-08-16 — One exact publisher abstract reduced the unassessed ledger to one. B92 adds a bounded non-AI multicriteria, Monte Carlo, and rank-stability method at the abstract layer. B91 remains content-unassessed; the 27-row ledger is 6 full / 20 abstract or official / 1 unassessed / 0 complete AI lifecycles.
-
2026-08-16 — Partial source lineage kept below a deployed build. The live DiscoverWater page differs from the sole published v1.2 page state, while eight of twelve dependency names and two of three canonical asset samples support partial shared lineage. The prototype demonstration and analytics hooks add no accepted-build, use, outcome, accessibility, cost, or AI receipt.
-
2026-08-16 — Final-three recovery found one public-delivery afterlife. B110 now has an official abstract, pinned framework and application source, and a same-named DiscoverWater interface reachable at KU. At that pass B92 and B91 remained content-unassessed. The 27-row ledger was then six full texts, 19 abstract or official-content rows, and two unassessed; no exact live-build join, use denominator, audience outcome, recheck, or AI lifecycle was admitted.
-
2026-08-16 — Residual DOI content recovery stopped at three unknowns. All 14 residual rows were investigated; eleven gained substantive primary content, including one full chapter, while three remain metadata-only. B102 adds an evaluation-to-redesign mechanism and B78/B38 add practice context, but no explicit GenAI role, accepted release, routine-use denominator, consequential outcome, or later affected-audience recheck.
-
2026-08-16 — Seven human evaluations kept below one deployed lifecycle. Nine high-signal DOI rows yield eight primary abstract or official-project surfaces and seven explicit participant/evaluator studies. InfoViP adds one NLP and unsupervised-learning production-intent near-miss; future-tense installation remains below accepted release, routine use, outcome, and later recheck. Fourteen DOI rows still lack substantive primary content.
-
2026-08-16 — Non-DOI recovery repaired authority before classifying evidence. Eight stable keys yielded three DOI recoveries, two year corrections, one rejected foreign PMID, four full texts, and one gated row. The strongest receipt is a non-AI cancer-diary embedded-use/usability comparator; zero complete AI lifecycles were added.
-
2026-08-16 — Human studies kept below deployment. Primary-text recovery for A208’s 27 no-abstract DOI rows puts five exact texts into custody: four empirical papers and one review. Controlled audience measurement, public- sector co-design/demo, and enterprise production intent remain three artifact-specific receipt lanes; none reaches accepted field release plus later affected-audience recheck. Twenty-two DOI rows remain unassessed.
-
2026-08-16 — Mature non-AI afterlife kept out of the AI lifecycle. Exact- DOI discovery covers 114 stable keys and 87 abstracts. AIDSVu is the sole five-category lifecycle signal, adding ten years of delivery, aggregate use, governance, and a later data release—but no AI role or affected-audience recheck. The remaining abstract and identifier gaps stay unclassified.
-
2026-08-16 — Artifact custody separated from audience delivery. A deterministic title-cue screen over 122 authoritative A208 keys surfaces one explicit GPT course study and one pinned 1,271-blob supplement. It yields zero audience-delivery or later-recheck chains; 121 title nonmatches remain not surfaced rather than excluded, and the lifecycle denominator stays null.
-
2026-08-16 — Longitudinal evaluation fills the development-decision middle, not the lifecycle. Lexara adds six developers using a deployed CVA evaluation toolkit for two weeks across 38 experiments, 57 newly authored cases, ten models, and six prompts. The combined twelve-receipt audit covers 14 named cases and still finds zero exact artifact carried through accepted delivery, consequential audience decision or calibrated trust, whole cost, and post-release same-lineage recheck.
-
2026-08-16 — Full six-domain crosswalk isolates five authority gaps. The 127 supplement positions now map to 122 distinct A208 review keys. Five exact-looking DOI candidates remain outside the review-controlled frame, one printed PMID resolves to an unrelated paper, and lifecycle classification remains unstarted and null.
-
2026-08-16 — First domain crosswalk finds a second authority gap. Eleven printed Education labels plus the B59 overlay resolve to twelve distinct A208 bibliography/DOI keys.
M. Lu (2020)has no A208 bibliography record; a plausible Crossref candidate stays outside the authoritative frame. The Education crosswalk remains 12 of 13 and lifecycle classification remains null. -
2026-08-16 — Education placeholder repaired; stable-key audit still open. A208’s post-selection Education paragraph and B59 bibliography uniquely reconstruct the published
?as Hernández-Calderón et al. (2023), DOI10.1093/iwc/iwac043. The gap ledger retains the broken source value, distinguishes the overlay from a publisher correction, and keeps lifecycle classification null until all rows have stable keys. -
2026-08-16 — Review inventory repair placed before lifecycle screening. A208’s supplement preserves domain totals of 127 but identifies 126 citations and one literal question-mark Education row. The gap ledger now keeps count, identity, and lifecycle classification separate; no placeholder becomes a negative finding.
-
2026-08-16 — A review corpus is not a lifecycle denominator. Added a 127-study decision-visualization review as a broad coverage map. Because the review does not specifically focus on GenAI and does not jointly extract the eight same-artifact receipts, the qualifying lifecycle count remains null, not zero; primary-paper classification is the next denominator-changing step.
-
2026-08-16 — Same endpoint, different retry burden, no cost winner. The VisCoder2 audit reconstructs 584 versus 714 conditional revisions over 888 tasks before both routes reach 732 execution passes. Missing comparable debug-resource telemetry and human or production acceptance keep the result at attempt topology; the total-cost ledger is now v7 with 0/12 accepted-cost comparisons.
-
2026-08-16 — Production recovery states separated by actor. Three held dashboard afterlives now resolve to two maintainer-verified restorations, zero affected-actor recoveries, zero authority transfers, and zero complete twelve-state rows; the gap ledger names the first valid upgrade receipt.
-
2026-08-15 — Reader outcomes split by AI role and lineage. A bounded ten-row audit now separates one direct controlled AI-created-chart reader effect, five adjacent human-outcome rows, three creator/co-design rows, and one model-reader-only row. Zero reaches accepted delivery plus later same- lineage reader recheck.
-
2026-08-15 — Governance split into two evidence rails. Three provider control contracts and two named-feature organizational-use reports are now shown as real but non-substitutable evidence, alongside two adjacent governance-process studies. Zero of seven held rows clears the nine-receipt same-deployment join.
-
2026-08-15 — Acquisition separated from adoption. Twelve named skill listings have public install signals, but collapse to ten parent repositories and eleven documented lineages. The seven-stage public ladder now keeps listing and acquisition separate from successful presence, invocation, retention, organizational acceptance, and outcomes; 0/12 named rows reaches the latter behavioral stages in held public evidence.
- 2026-08-15 — Eleven-stage system anatomy added. Six primary authoring or analysis systems all separate generation, execution, and some critique route; four expose meaningful human control; zero reaches accepted delivery, intended-reader outcome, or later maintenance. The public map assigns information and authority without selecting one universal agent topology.
- 2026-08-15 — The capability-to-practice bridge is short, not empty. Ten primary cases yield three controlled partial joins—to analyst use, analyst correction, and reader harm—but zero exact-version accepted-delivery bridge and zero complete episode. The gap ledger retains the first missing receipt rather than composing a lifecycle across studies.
- 2026-08-15 — Fixed-budget near miss added. Selective TTS matches declared LLM-call and separate output-token budgets inside one visual-insights pipeline, but its released accounting omits repair and errored-worker cost and its candidates are not contract-accepted artifacts. The maintenance and cost row now distinguishes declared partial budget, observed route-wide use, and accepted-output cost; the complete result is 0/11.
- 2026-08-15 — Six-receipt lifecycle null made explicit. Seven leading cases were audited across repeated representative use, immutable tested build, exact model, versioned release, later event, and post-change recheck. Zero clears the full join; the null is search-scoped and retains every partial lane and reopen condition.
- 2026-08-15 — Longitudinal co-design added without lifecycle promotion. Graphy follows three blind co-designers through 12 sessions, four workshops, eight months, and implemented interaction changes. Its paper disclaims formal evaluation, and its seven-commit public repository has no workshop binding, release, or representative recheck; zero new complete episodes were added.
- 2026-08-15 — Accepted-artifact denominator added. ChartAgent’s accuracy/tool-call sweep adds one partial cost proxy, but zero of nine held fragments has equivalent route cost plus frozen accepted outcomes. Ledger v5 retains failures in the numerator and rejects benchmark accuracy, tool calls, and per-attempt averages as accepted-cost substitutes.
- 2026-08-15 — Study surface found; tested build still unknown. MAIDR’s
public legacy code now supplies a dedicated AI-study surface, a
v2.10.0version floor, and real later maintenance. The abstract-level eight-person result has no exact tested-build or model join, the later changes have no representative recheck, and the TypeScript rewrite is a new validation target. - 2026-08-15 — Four near-misses, zero new lifecycle episodes. Added a seven-key horizontal admission rule, representative pre-AI MAIDR accessibility evidence, the current AI rewrite boundary, and a described-but-removed human study. Useful evidence stays in its lane; project history does not become an artifact afterlife by aggregation.
- 2026-08-15 — Five validity threats block configuration promotion. Added CHART-6 and a high-level human/model interpretation study, separated task success from human-like errors and reader experience, forbade benchmark lift, and kept the qualifying current Vizier configuration-run count at zero.
- 2026-08-15 — Machine totals need phase and topology. Added adjacent ProMCP profiling and a pinned release audit. The total-cost ledger now keeps cold start, discovery, execution phase, effective topology and runtime state visible; no visualization comparison, reproduced run, or cost winner was added.
- 2026-08-15 — Contribution is a state, not a handoff conclusion. Added accepted non-owner AI-assisted work in Prism and OpenClaw, Prism’s repeat contribution, and OpenClaw’s open repair proposal. The evaluator must keep contribution, authority, and recovery separate; no complete lifecycle comparison exists.
- 2026-08-15 — Repaired release separated from independent recovery. Added Prism as a third same-artifact dashboard afterlife and the first held multi-attempt delivery chain. It reaches corrected release and maintainer runtime verification, not independent final acceptance; no complete cost or lifecycle comparison exists.
- 2026-08-15 — Equal rounds are not equal resources. Added METAL as the seventh specialist-cost fragment. Its routes match at five candidates or recurrences but not call topology, and the paper and released logs lack the observed per-arm usage needed for an equal-budget claim. The ledger now requires planned and observed receipts across twelve resource dimensions; no run or cost winner exists.
- 2026-08-15 — Total cost starts before inference. Added a sixth fragment, quarantined its exact figures after four source-integrity failures, and expanded the protocol to eleven lanes, five cost classes, declared amortization, and a source preflight. No route ran and no winner exists.
- 2026-08-15 — One production afterlife, no complete lifecycle. Added KubeStellar’s self-reported same-dashboard regression, two-step repair, restored deploy, and prevention chain plus PM4Py-UCM’s measured refinement effort; preserved the missing whole cost, comparison, handoff, accessible reader, decision, and trust lanes.
- 2026-08-15 — Total workflow cost gets a denominator, not a winner. Audited five primary cost fragments and added a same-task, equal-budget, ten-lane ledger through accepted delivery, reader use, and a later change. The held record supports no summed total or specialist-versus-general ranking.
- 2026-08-15 — Longitudinal learning gap made runnable, not resolved. Added the bounded search null and a three-arm, five-wave creator design with six-month unassisted transfer, a separate blinded reader stage, ten outcome lanes, and ten unfulfilled human/experimental release gates.
- 2026-08-15 — Typed interaction plan added; forecast still unchanged. Added ViviDoc’s State–Render–Transition–Constraint mechanism, same-pipeline ablation, creator/rater evidence, and reader/lifecycle limits; disclosed that it predates the forecast freeze and resolves none of the seven tests.
- 2026-08-15 — Editor calibration task prepared, not performed. Added ten stimulus-by-lens baselines, ten synthetic quality vignettes with two swapped repeats, three response states, 14 rejected adverse records, and seven audience decisions. No editor, appointment, calibration, tie, call, contact, ranking, or result exists.
- 2026-08-15 — Seven R4 receipts defined; all real answers pending. Added three per-slot provider receipts, editor acceptance and calibration, a capped 30-primary-call budget, exact-set release, 20 admissible paths, 16 rejected adverse states, and seven audience decisions. No provider, editor, budget, release, call, contact, ranking, or result exists.
- 2026-08-15 — Two R4 prompts, two schemas, zero model answers. Added separate integrity and readability prompts, three honest raw-response states, twelve rejected aggregate/lens/identity/evidence/state failures, and seven audience decisions. Exact models, editor calibration, budget, release, calls, contacts, rankings, and results remain absent.
- 2026-08-15 — Five R4 inputs frozen; no model run. Added five opaque-ID SVGs, a 19-file model-input allowlist, nine source-evidence files, five evaluator-only truth records, adverse leakage/hash checks, corrected visual-inspection findings, and seven audience decisions. Model slots, editor, tie, prompts, budget, release, calls, and results remain absent.
- 2026-08-15 — R4 is preregistered, not run. Added five candidate contexts, three empty exact-model slots, separate integrity/readability lenses, a human-editor and tie-rule gate, seven audience translations, and a pilot boundary. No stimulus, ranking, call, contact, or divergence exists.
- 2026-08-15 — Outcomes stay separated instead of becoming one score.
Added two mechanical tiers, six exact nonmechanical lanes, authority binding,
90 preserved synthetic
not-runstates, aggregate-score rejection, and seven audience translations. No artifact, arm winner, human result, or efficacy claim is implied. - 2026-08-15 — Six gates become an attributable decision register. Added 17 admissible paths across four authority roles, fail-closed response and final-release rules, and seven audience translations. All six selections and release remain pending; no provider fact, approval, spend, contact, call, or efficacy result is implied.
- 2026-08-15 — Acceptance authority stays typed and human. Added six
nonmechanical authority lanes, five fixture-specific role maps, a 30-state
mechanical-only
not-runpolicy, W3C’s conformance/user-evaluation boundary, seven audience decisions, and the research-owner gate. No person, model call, human verdict, or scope selection is implied. - 2026-08-15 — One-layer candidate separates mechanism from skill bundle. Added one post-artifact evidence-reconciliation candidate, its five-part causal-unit test, five-context failure target, seven audience decisions, and the owner-ratification boundary. No selection, call, or effect is claimed.
- 2026-08-15 — Credential-free runner separates task from evaluator. Added five source checks, 15 isolated fixture-arm envelopes, an explicit model- input/evaluator-only boundary, external artifact/render/trace/check hashing, fail-closed receipt identity, six residual blockers, and seven audience translations. No model, artifact, browser, human, or efficacy result is implied.
- 2026-08-15 — Current-harness manifest preflight. Added the exact fields that can be frozen before a call, an 18-field equivalent-receipt contract, seven named blockers, zero-call/no-spend evidence, and audience-specific rules for interpreting a workflow comparison.
- 2026-08-15 — Interactive-application source packet. Added the fifth of five E0 source packets, including exact canonical state, synchronized surfaces, replay, history, reset, export, keyboard, narrow delivery, stale-selection recovery, and prepared-change gates with explicit browser, maintenance, accessibility, and human-evidence limits.
- 2026-08-15 — Specialized uncertainty source packet. Added the fourth of five E0 source packets, including exact official estimate, margin-of-error, geography, pairwise-test, misleading-rank, refusal, render, and expert-review gates with explicit cross-domain and human-evidence limits.
- 2026-08-15 — Governed-BI source packet. Added the third of five E0 source packets, including exact semantic, vintage, permission, verified-query, refusal, fanout, and incident-recovery gates with explicit production and human-evidence limits.
- 2026-08-15 — Exploratory source packet. Added the second of five E0 source packets, including exact pooled/group checks, revision receipts, audience-specific use boundaries, and an explicit familiar-data limit.
- 2026-08-15 — Shared-baseline readiness gate. Added the reconciled historical custody boundary, the initial context-readiness state, and the equivalent-receipt requirement for bare, normal-harness, and one-layer arms.
- 2026-08-14 — Initial public edition. Established the contextual goals typology, literature synthesis, technique history, benchmark timeline, convergence and divergence findings, demonstrated negatives, and research gap ledger.