A moving model-and-harness baseline

Models are improving, so generic instructions expire faster than local context and checks.

There is no clean visualization-specific curve covering August 2023 through August 2026. Benchmarks, judges, models, prompts, and test-time budgets changed. The literature supports a direction and several measured slices—not a universal annual improvement rate or a defensible straight-line forecast.

How to read the evidence

Improvement is large in fidelity and breadth, smaller in basic execution, and uneven by task.

The best comparable results come from papers that evaluate multiple model generations on the same benchmark. Even those slices are not a pure measure of training progress: provider, model size, vision stack, inference policy, and evaluator can all change together.

The widening task envelope is just as important as the score movement. In 2023, a demanding benchmark asked a model to reproduce one scientific plot. By 2026, DashArena asks a general model to construct a multi-view interactive dashboard and declare a replayable analysis path. Broader capability does not mean dependable behavior.

How to read the paired values. Each row has its own labeled scale. Compare the open earlier point with the filled later point inside that row; horizontal positions are not comparable across different metrics.

Plot2Code · direct, single passSep 2023 → Jun 2025

Twenty-one months of movement on one scientific chart-to-code benchmark

Code executes percent of examples○ 84.1%● 87.9%
0%+3.8 percentage points100%
Text match percent of text matched○ 48.5%● 71.7%
0%+23.2 percentage points100%
Rendered quality score from 0 to 10○ 5.45 / 10● 7.65 / 10
0+2.20 scale points10

An anonymous 2026 preprint re-reported the same Python/Matplotlib subset and normalized the old and new model results over the full test set. Useful trend evidence; not peer-reviewed longitudinal proof.

CharTide · same-paper comparisonGPT-4o → GPT-5

A later general model improves quality more than execution

ChartMimic high-level score from 0 to 100○ 87.7● 94.7
0+7.0 scale points100
Plot2Code text match percent of text matched○ 52.6%● 61.9%
0%+9.3 percentage points100%
ChartX score score from 0 to 5○ 2.61 / 5● 3.59 / 5
0+0.98 scale points5

ACL 2026. The chart-specialized CharTide models also matched or exceeded general frontier models, showing that better baselines can come from domain training as well as scale.

Targeted feedback helped

Text2Vis: 26% direct → 42% with one answer-and-code review

Three examples left GPT-4o at 26%; retrieval plus examples reached 31%. One structured answer-and-code feedback round reached 42%. Adding visual feedback improved visual subscores but slightly reduced final pass to 41%.

Interpretation: add the missing evidence channel; do not equate another prompt or model call with progress.

More instruction sometimes hurt

Plot2Code: stricter requirements traded execution for resemblance

Detailed conditional instructions generally improved similarity while lowering code pass rates; Gemini Pro fell from 68.2% to 55.3%. Chain-of-thought and Plan-and-Solve showed no clear advantage over the default prompt.

Interpretation: prose scaffolding can collide with the task, budget, or model rather than simply add knowledge.

Breadth outran reliability

DashArena: full dashboard generation is now plausible, not dependable

Current models can generate an interactive dashboard and intended-use trajectory from an open task. The strongest aggregate preference result was competitive with the human baseline, but no model cleared 86% rendering or 74% replay.

Interpretation: the frontier moved from “can it attempt this?” to “where and how does it silently fail?”

State is part of correctness

Dashboard2Code: complex interaction remains much harder

The best configuration scored 79.4 overall and 64.2 on the hardest interaction level across 180 Plotly Dash applications. DOM access raised one model’s successful UI exploration from 20.61% to 64.58%.

Interpretation: a dashboard can respond while hidden state or a transformation is wrong. Fixed 1920×1080 desktop, no animation or popups.

Professional reading is a separate test

Chartography: the best configuration passed 45% of deliberately hard tasks

One hundred practitioner-authored tasks span twelve domains and have three expert verifiers each. More reasoning usually helped, but often elaborated an initial visual misread.

Interpretation: this adversarial slice is not a failure rate for ordinary charts; it shows why basic chart-QA success is not a professional release gate.

Preference did not imply accuracy

BLV learners preferred tactile + text + chat, with no measured accuracy lift

Eleven of twelve participants preferred the multimodal condition and described a better spatial mental model. The chart-understanding accuracy comparison did not improve.

Interpretation: measure modality, mental model, preference, and comprehension separately.

Practitioner + editor

Keep the answer key out of the task

Inspect the model-input manifest. Expected values, known failures, rubrics, and editorial judgment belong outside the generation context.

Decision: whether the comparison tests work or evaluator leakage.

Product + analytics

Require an exportable pilot directory

Retain source, role, semantic context, artifact, renders, checks, failures, human state, and cost per task and arm—not a screenshot plus score.

Decision: whether pilot evidence can be audited.

Research

Stage first, invoke second

Predeclare visibility, isolate every attempt, hash outputs outside the arm, and report missing receipts as missing before comparing outcomes.

Decision: whether every arm faced the same evaluable task.

Learning + access

Checks are not people

A check can establish that files, labels, viewports, or states exist. Comprehension, transfer, keyboard experience, and assistive-technology use remain not-run.

Decision: which acceptance layer still needs a person.

Practitioner + editor

Name the one checkpoint

If form, context, tools, retries, and critique changed too, report the workflow bundle. Attribute one mechanism only when it alone moved.

Decision: whether a local practice deserves causal credit.

Product + analytics

Fix the insertion and target

State where the control runs, which silent number or state divergence it should catch, what stays equal, and how false alarms and review cost are counted.

Decision: whether the pilot tests a feature or a package.

Research + engineering

Register the delta, not the package

One mechanism, one output schema, equal budgets, evaluator truth outside model input, and a separately versioned provider adapter.

Decision: whether any outcome can be attributed.

Learning + access + sponsor

Prepared is not selected or proven

A discrepancy check cannot supply human acceptance, and a passing protocol verifier cannot supply owner approval or an effect estimate.

Decision: which authority and evidence gate comes next.

Checks

Sign facts, not judgments

Deterministic evaluators own exact source, transform, artifact, state, browser, and receipt findings—not usefulness or comprehension.

Decision: whether the result is mechanical-only.

Context + domain

Name the decision owner

Editors, analysts, finance and governance owners, survey-methods experts, and operations managers answer different questions.

Decision: whose rationale settles the stated scope.

Accessibility

Conformance is not experience

Knowledgeable evaluation and relevant disabled-user task evidence complement each other; neither can be replaced by a scanner.

Decision: which accessibility claim was actually tested.

People + afterlife

Contribution is not authority

Prism and OpenClaw preserve accepted second-person AI-assisted changes; Prism has a repeat contributor, and OpenClaw has an open repair proposal. Keep accepted change, repeat contribution, maintenance authority, and independent recovery separate.

Decision: which claim must stop at the mechanical layer.

Provider

Attest facts, not permission

A served snapshot or exported harness receipt establishes what is available. It does not authorize spend, amend the protocol, or select a product layer.

Decision: which provider fact is actually in custody.

Owners

Answer one gate at a time

Research and Vizier owners choose only within their named lane. A passing verifier, model label, or credential name is never an inferred yes.

Decision: whose authority changes this exact field.

Protocol + spend

Make the alternative explicit

Dropping the bare arm needs a versioned amendment and narrower claims. Every paid or no-cost route still needs call, token, and stop caps.

Decision: what changed and what the run may now claim.

Release

Release the resolved set

Partial answers stay blocked. Six advancing receipts still need one final owner signature over their digest, run plan, maximum calls, and stop condition.

Decision: whether this exact run may make its first call.

Missingness

Missing stays missing

An empty mechanical tier or any not-run check is incomplete. A human not-run is never encoded as zero, tie, or pass.

Decision: which evidence was actually produced.

Authority

One role signs one lane

Context, domain, conformance, disabled-user, intended-user, and maintainer authority are typed and cannot substitute for one another.

Decision: whose evidence supports this exact outcome.

Trade-offs

Keep the outcome vector

A gain in one lane cannot erase a failure, regression, or unrun result elsewhere. Compare arms lane by lane.

Decision: what improved, regressed, or remains unknown.

Cost

Put cost beside outcomes

Calls, spend, latency, repairs, and human time remain visible rather than disappearing inside a quality grade.

Decision: whether a bounded gain is worth its full cost.