Models are improving, so generic instructions expire faster than local context and checks.
There is no clean visualization-specific curve covering August 2023 through August 2026. Benchmarks, judges, models, prompts, and test-time budgets changed. The literature supports a direction and several measured slices—not a universal annual improvement rate or a defensible straight-line forecast.
How to read the evidence
Improvement is large in fidelity and breadth, smaller in basic execution, and uneven by task.
The best comparable results come from papers that evaluate multiple model generations on the same benchmark. Even those slices are not a pure measure of training progress: provider, model size, vision stack, inference policy, and evaluator can all change together.
The widening task envelope is just as important as the score movement. In 2023, a demanding benchmark asked a model to reproduce one scientific plot. By 2026, DashArena asks a general model to construct a multi-view interactive dashboard and declare a replayable analysis path. Broader capability does not mean dependable behavior.
How to read the paired values. Each row has its own labeled scale. Compare the open earlier point with the filled later point inside that row; horizontal positions are not comparable across different metrics.
Plot2Code · direct, single passSep 2023 → Jun 2025
Twenty-one months of movement on one scientific chart-to-code benchmark
Code executes percent of examples○ 84.1%→● 87.9%
Text match percent of text matched○ 48.5%→● 71.7%
Rendered quality score from 0 to 10○ 5.45 / 10→● 7.65 / 10
An anonymous 2026 preprint re-reported the same Python/Matplotlib subset and normalized the old and new model results over the full test set. Useful trend evidence; not peer-reviewed longitudinal proof.
CharTide · same-paper comparisonGPT-4o → GPT-5
A later general model improves quality more than execution
ChartMimic high-level score from 0 to 100○ 87.7→● 94.7
Plot2Code text match percent of text matched○ 52.6%→● 61.9%
ChartX score score from 0 to 5○ 2.61 / 5→● 3.59 / 5
ACL 2026. The chart-specialized CharTide models also matched or exceeded general frontier models, showing that better baselines can come from domain training as well as scale.
Targeted feedback helped
Text2Vis: 26% direct → 42% with one answer-and-code review
Three examples left GPT-4o at 26%; retrieval plus examples reached 31%. One structured answer-and-code feedback round reached 42%. Adding visual feedback improved visual subscores but slightly reduced final pass to 41%.
Interpretation: add the missing evidence channel; do not equate another prompt or model call with progress.
More instruction sometimes hurt
Plot2Code: stricter requirements traded execution for resemblance
Detailed conditional instructions generally improved similarity while lowering code pass rates; Gemini Pro fell from 68.2% to 55.3%. Chain-of-thought and Plan-and-Solve showed no clear advantage over the default prompt.
Interpretation: prose scaffolding can collide with the task, budget, or model rather than simply add knowledge.
Breadth outran reliability
DashArena: full dashboard generation is now plausible, not dependable
Current models can generate an interactive dashboard and intended-use trajectory from an open task. The strongest aggregate preference result was competitive with the human baseline, but no model cleared 86% rendering or 74% replay.
Interpretation: the frontier moved from “can it attempt this?” to “where and how does it silently fail?”
State is part of correctness
Dashboard2Code: complex interaction remains much harder
The best configuration scored 79.4 overall and 64.2 on the hardest interaction level across 180 Plotly Dash applications. DOM access raised one model’s successful UI exploration from 20.61% to 64.58%.
Interpretation: a dashboard can respond while hidden state or a transformation is wrong. Fixed 1920×1080 desktop, no animation or popups.
Professional reading is a separate test
Chartography: the best configuration passed 45% of deliberately hard tasks
One hundred practitioner-authored tasks span twelve domains and have three expert verifiers each. More reasoning usually helped, but often elaborated an initial visual misread.
Interpretation: this adversarial slice is not a failure rate for ordinary charts; it shows why basic chart-QA success is not a professional release gate.
Preference did not imply accuracy
BLV learners preferred tactile + text + chat, with no measured accuracy lift
Eleven of twelve participants preferred the multimodal condition and described a better spatial mental model. The chart-understanding accuracy comparison did not improve.
Interpretation: measure modality, mental model, preference, and comprehension separately.
Practitioner + editor
Keep the answer key out of the task
Inspect the model-input manifest. Expected values, known failures, rubrics, and editorial judgment belong outside the generation context.
Decision: whether the comparison tests work or evaluator leakage.
Product + analytics
Require an exportable pilot directory
Retain source, role, semantic context, artifact, renders, checks, failures, human state, and cost per task and arm—not a screenshot plus score.
Decision: whether pilot evidence can be audited.
Research
Stage first, invoke second
Predeclare visibility, isolate every attempt, hash outputs outside the arm, and report missing receipts as missing before comparing outcomes.
Decision: whether every arm faced the same evaluable task.
Learning + access
Checks are not people
A check can establish that files, labels, viewports, or states exist. Comprehension, transfer, keyboard experience, and assistive-technology use remain not-run.
Decision: which acceptance layer still needs a person.
Practitioner + editor
Name the one checkpoint
If form, context, tools, retries, and critique changed too, report the workflow bundle. Attribute one mechanism only when it alone moved.
Decision: whether a local practice deserves causal credit.
Product + analytics
Fix the insertion and target
State where the control runs, which silent number or state divergence it should catch, what stays equal, and how false alarms and review cost are counted.
Decision: whether the pilot tests a feature or a package.
Research + engineering
Register the delta, not the package
One mechanism, one output schema, equal budgets, evaluator truth outside model input, and a separately versioned provider adapter.
Decision: whether any outcome can be attributed.
Learning + access + sponsor
Prepared is not selected or proven
A discrepancy check cannot supply human acceptance, and a passing protocol verifier cannot supply owner approval or an effect estimate.
Decision: which authority and evidence gate comes next.
Checks
Sign facts, not judgments
Deterministic evaluators own exact source, transform, artifact, state, browser, and receipt findings—not usefulness or comprehension.
Decision: whether the result is mechanical-only.
Context + domain
Name the decision owner
Editors, analysts, finance and governance owners, survey-methods experts, and operations managers answer different questions.
Decision: whose rationale settles the stated scope.
Accessibility
Conformance is not experience
Knowledgeable evaluation and relevant disabled-user task evidence complement each other; neither can be replaced by a scanner.
Decision: which accessibility claim was actually tested.
People + afterlife
Contribution is not authority
Prism and OpenClaw preserve accepted second-person AI-assisted changes; Prism has a repeat contributor, and OpenClaw has an open repair proposal. Keep accepted change, repeat contribution, maintenance authority, and independent recovery separate.
Decision: which claim must stop at the mechanical layer.
Provider
Attest facts, not permission
A served snapshot or exported harness receipt establishes what is available. It does not authorize spend, amend the protocol, or select a product layer.
Decision: which provider fact is actually in custody.
Owners
Answer one gate at a time
Research and Vizier owners choose only within their named lane. A passing verifier, model label, or credential name is never an inferred yes.
Decision: whose authority changes this exact field.
Protocol + spend
Make the alternative explicit
Dropping the bare arm needs a versioned amendment and narrower claims. Every paid or no-cost route still needs call, token, and stop caps.
Decision: what changed and what the run may now claim.
Release
Release the resolved set
Partial answers stay blocked. Six advancing receipts still need one final owner signature over their digest, run plan, maximum calls, and stop condition.
Decision: whether this exact run may make its first call.
Missingness
Missing stays missing
An empty mechanical tier or any not-run check is incomplete. A human not-run is never encoded as zero, tie, or pass.
Decision: which evidence was actually produced.
Authority
One role signs one lane
Context, domain, conformance, disabled-user, intended-user, and maintainer authority are typed and cannot substitute for one another.
Decision: whose evidence supports this exact outcome.
Trade-offs
Keep the outcome vector
A gain in one lane cannot erase a failure, regression, or unrun result elsewhere. Compare arms lane by lane.
Decision: what improved, regressed, or remains unknown.
Cost
Put cost beside outcomes
Calls, spend, latency, repairs, and human time remain visible rather than disappearing inside a quality grade.
Decision: whether a bounded gain is worth its full cost.