Research source document · Evidence reviewed through August 18, 2026

Specialized vision models for data visualization: what they can—and cannot—verify

Status: research snapshot, evidence cut 2026-08-16. Recheck by 2026-11-16, or earlier after a major general-model release, independent replication, or specialist-assisted reader study.

Companion: The state of AI-assisted data visualization research.

This report examines specialized vision models and systems that might help an AI understand or critique a data visualization. It covers chart question answering, chart-to-table extraction, document OCR, chart-element grounding, misleading-chart detection, perceptual quality judgment, and verifier or repair loops. It asks where specialization adds something beyond a strong general multimodal model—and where the available evidence does not support that claim.

Executive assessment

There is no specialist model that can serve as a trustworthy final judge of a data visualization. There are useful specialists for different parts of the problem.

A chart parser can recover values without knowing whether the chart supports its headline. A grounding model can locate a line or legend without verifying the source data. A misleading-chart detector can identify a truncated axis while missing a false annotation. A learned quality critic can reproduce expert ratings of composition and readability without knowing whether readers understand the result. These are different jobs with different evidence.

The findings support six conclusions:

  1. “Specialized” is not a quality ranking. Chart-specific models that led older benchmarks performed poorly when ChartQAPro introduced realistic, harder, and out-of-distribution charts. On a separate 2026 extraction benchmark, a frontier general model substantially outscored DePlot and TinyChart.
  2. Specialized tools can still add large gains to a general reasoner. In ChartAgent, a chart-specific tool system improved the same GPT-4o model by 26.7 percentage points on unannotated numeric questions; chart-specific tools beat generic image tools by 30.0 points overall.
  3. The most direct specialist critic is promising but not a final quality gate. VisJudge predicts its expert-adjudicated visualization ratings better than the tested general models. Its chart-revision gains were scored by another model, not by readers or human experts, and it cannot verify source fidelity from an image alone.
  4. Better perception is useful evidence, not a complete critique. New OCR, structured-graphics, and grounding models can recover text, layout, values, and marks. None of the captured evaluations shows that a stronger parser by itself produces better end-to-end judgments about communication, integrity, or reader outcomes.
  5. The defensible architecture is composed and context-dependent. Start from source data and chart specification, run deterministic checks, add the smallest specialist that addresses a measured perception failure, retain a current general model for broad reasoning, and preserve human acceptance. When comprehension, decisions, mobile use, or accessibility matter, test the delivered experience with the intended readers.
  6. Machine totals need phase and topology. Adjacent ProMCP evidence shows measured burden moving between planning/schema injection and answer synthesis as client topology changes. Preserve cold start, discovery, planning, tool/result transfer, context update, and synthesis separately; this is a measurement rule, not a visualization cost winner.

The specialized-vision gap in the broader field review should therefore move from open to partly answered. The field now has credible evidence for several component roles and clear negative findings. It still lacks a complete routing study, independent replications, end-to-end cost evidence, and human reader outcomes that isolate the specialist’s contribution.

“Vision for charts” contains six different problems

Role Input and output What success means What it cannot establish alone
Chart QA and reasoning chart image + question → answer correct answer on a declared chart distribution source fidelity, overall quality, reader comprehension
Chart parsing chart image → table, values, or structured graphic recovered labels, series, geometry, and numbers whether the visual encoding or message is appropriate
Document OCR and layout page or screenshot → text, regions, tables, formulas, reading order accurate document structure under realistic degradation chart semantics, insight, integrity, or aesthetic quality
Grounding and segmentation image + referring phrase → point, box, or mask the requested mark or region is localized whether the mark is true, legible, or meaningful
Integrity and quality critique chart image, sometimes with context → defect labels or ratings agreement with a defined taxonomy or human rating protocol dimensions excluded from the rubric or unavailable source evidence
Verifier and repair component repeated outputs, intermediate evidence, or candidate revision → score, uncertainty, or edit errors become detectable and repair improves a held-out outcome correctness outside the tested loop or after delivery

This typology is the organizing principle for the rest of the report. Model size, release date, and the word “chart” in a benchmark name are secondary to the role being measured.

A short history: progress and harder tests arrived together

The apparent story changes depending on whether the timeline tracks model scores or the questions researchers learned to ask.

Date Work What changed What later evidence qualified
Dec. 2022–2023 DePlot Converted a chart image to a linearized table so a language model could reason over values; DePlot + LLM reached 67.6 on ChartQA’s human split versus 38.2 for the prior MatCha system. Tables discard color, orientation, geometry, and other visual evidence; DePlot scored 23.06 RMSF1 on a harder 2026 extraction test.
Jan.–Apr. 2024 ChartAssistant, TinyChart Added staged chart-table alignment, high-resolution visual token merging, and program-of-thought arithmetic in compact models. TinyChart’s program execution raised calculative accuracy from 56.64 to 78.98 on its older evaluation. Both families transferred poorly to ChartQAPro; older benchmark leadership is not stable evidence of realistic chart understanding.
Jun.–Sep. 2024 CharXiv, ChartGemma, EvoChart Scientific charts became a distinct test; synthetic instructions and self-training expanded chart-specialist training data. Generated training and author benchmarks cover bounded chart distributions. ChartGemma scored 9.80 with chain-of-thought on ChartQAPro.
Apr.–May 2025 ChartQAPro, CHART-6 Replaced templated questions with human-written, realistic, conversational, hypothetical, fact-checking, and unanswerable tasks; compared model errors with six human visualization-literacy tests. Best reported ChartQAPro result was 55.81 versus an 85.02 human baseline; no single benchmark represents all visualization literacy.
Aug.–Oct. 2025 Misviz, VisJudge Shifted from answering questions to detecting misleading designs and predicting expert-adjudicated quality across single charts, multiple charts, and dashboards. Real and synthetic misleading-chart rankings differ. Quality prediction still lacks raw-data fidelity and reader-outcome evidence.
Jan.–Mar. 2026 PaddleOCR-VL-1.5, dots.mocr Sub-billion to 3B models reported strong document parsing; structured graphic and SVG recovery became an explicit output. Document parsing accuracy is an upstream capability, not a visualization-quality score. Some results rely on author benchmarks or model judges.
Apr.–May 2026 ExChart, ChartAgent, ChartREG++, self-ensembling extraction Systems added chart-calibrated human correction, specialist tools, mark-level grounding, repeated sampling, and uncertainty. Tools fail, candidate masks constrain results, inference grows 4–16×, and uncertainty remains only modestly correlated with correctness.
Jul.–Aug. 2026 HunyuanOCR-1.5, VisDeception, Dashboard2Code, Touching or Chatting, scientific visualization literacy, accessible chart descriptions, CURV, Chartography Lightweight document parsing improved; deception mitigation, stateful dashboard reconstruction, BLV learning, professional chart reading, newer frontier-model literacy, context ablations, and explicit grounded-reasoning curricula arrived. Multi-stage mitigation can hurt some models. Stateful evaluation remains fixed-desktop; the BLV study finds preference and spatial-model benefit without an accuracy lift; several works are recent preprints.

The field has improved over 36 months, but the timeline does not support one smooth capability curve. As models improved, benchmarks expanded from clean single charts and short answers to scientific figures, dashboards, unanswerable questions, deceptive designs, element localization, and accessible descriptions. A score from 2023 and a score from 2026 often answer different questions.

What the strongest same-test comparisons show

The table below deliberately avoids cross-benchmark ranking. Every comparison within a row uses the benchmark, metric, and condition reported by the same study.

Problem and benchmark Specialist or intervention General or base comparison Supported reading
Realistic chart QA — ChartQAPro, chain-of-thought accuracy ChartGemma 9.80; TinyChart 14.24 Claude Sonnet 3.5 55.81 Historical chart specialists did not transfer to this harder distribution.
OOD chart-to-table — WB-ChartExtract, single-pass RMSF1 DePlot 23.06; TinyChart 28.61 Gemini 2.5 Pro 87.83; Claude Opus 4.6 60.99; GPT-5.1 51.26 General frontier models can dominate specialists on a new extraction distribution.
Chart QA with tools — ChartBench overall accuracy ChartAgent 71.39 GPT-4o 54.53; Phi-3-Vision 55.32 A general orchestrator with chart-specific tools can outperform standalone systems.
Chart QA tool ablation — ChartAgent overall accuracy chart-specific image + analysis tools: +30.0 points over generic image tools same orchestrator and benchmark The gain comes from task-relevant executable tools, not simply adding an agent loop.
Quality rating — VisJudge held-out MAE / correlation VisJudge-7B .421 / .687 GPT-5 .553 / .428; GPT-4o .610 / .482; Claude 4 Sonnet .622 / .465 A small task-specific critic better matches this expert-adjudicated rubric.
Misleading-chart detection — Misviz real-world F1 image + extracted-axis classifier 60.9; linter + extracted axis 36.1 GPT-o3 80.0; GPT-4.1 79.6 General MLLMs lead on heterogeneous real-world visualizations.
Misleading-chart detection — Misviz synthetic F1 linter + predicted axis 67.9; image + axis classifier 65.2 GPT-o3 66.9 Rules and classifiers remain competitive on the controlled distribution they encode.
Chart grounding — ChartREG++ point/box F1 best reported point F1 remained below 50; chart-specific mask candidates + alignment reached 52.55/49.85 mask F1 on two sets SAM 3 zero-shot candidates + alignment reached 25.01/21.84 Fine chart marks remain hard; chart-aware candidates help, but this is not a fair zero-shot model contest.
Deception mitigation — VisDeception error rate structured-metadata multi-stage system reduced Gemini 2.5 Pro .099→.024 and GPT-4o .502→.255 same models without mitigation Explicit decomposition can help substantially, but not universally. Gemini 2.5 Flash worsened .109→.135.
Explicit grounded-reasoning curriculum — public benchmarks CURV-7B improved its Qwen base: ChartQAPro 29.77→32.80 and CharXiv 32.50→36.70 same base checkpoint Training explicit grounding improves the selected base, but the absolute hard-benchmark gap remains and the preprint lacks external replication.

These results settle one question and open another. There is no class-wide specialist advantage. There are strong component-level interventions. The system-design problem is to route each only where its information or executable capability adds to a current baseline.

Chart reasoning: compact specialists improved the old tasks, not the whole problem

DePlot established a durable idea: do not ask one model to perceive the chart and calculate the answer in a single hidden step. It first converts the image to a table, then lets a language model answer over that table. The decomposition made values inspectable and produced a large ChartQA gain. It also created an information bottleneck. A table preserves values and labels; it does not preserve color, line style, spatial emphasis, annotations, or interactive state.

ChartAssistant and TinyChart developed the same general direction differently. ChartAssistant first aligned charts with tables, then performed multi-task instruction tuning. TinyChart compressed high-resolution visual tokens and emitted Python for arithmetic questions. Both show that explicit alignment and executable calculation can help. They also show why an older benchmark score should not be mistaken for a lasting model choice.

ChartGemma used a 3B PaliGemma model and 122,857 charts, with Gemini-generated instructions. An audit of 100 generated instructions found 82% accurate and 8% partially accurate; only two volunteers performed the human evaluation. It reported strong results on older benchmarks, then scored 9.80 with chain-of-thought on ChartQAPro. This is not evidence that ChartGemma was poorly designed. It is evidence that benchmark fit was much narrower than the label “chart understanding.”

EvoChart scaled synthetic self-training to 1.6 million question-answer pairs. Its 4B model reached 54.2 on the authors’ EvoChart-QA test versus 49.8 for GPT-4o. That is a legitimate same-test result, with two cautions: the paper gives inconsistent 625/650 counts for the benchmark charts, and the test covers four common chart types and basic comprehension. On modified familiar charts, GPT-4o fell from 85.7 to 45.8 and TinyChart from 83.6 to 45.8. Small visual changes exposed brittle pattern matching in both general and specialist systems.

ChartQAPro is important because it changes the task. Its 1,341 charts come from 157 sources and include dashboards, infographics, and multiple charts. Its 1,948 human-written and verified questions include conversations, hypotheticals, fact checking, and unanswerable cases. The best reported result, Claude Sonnet 3.5 with chain-of-thought, was 55.81 versus an 85.02 human baseline. The human estimate comes from one expert graduate student on 50 questions per category, so it is not a population norm. The benchmark is still static and excludes interactive dashboards. It is a better transfer test, not the final definition of chart literacy.

The August 2026 CURV preprint trains explicit visual grounding before harder reasoning. CURV-7B improves its Qwen 2.5-VL-7B base across ChartQA, ChartQAPro, CharXiv, ChartMuseum, MathVista, and MMMU-Pro. That breadth makes it worth watching. It remains a new, unreplicated preprint; some answer evaluation uses GPT-4.1-mini, and level-three multi-chart performance remains at or below 42.54.

Current reading: chart-tuned training is most credible when it improves a current base across several held-out distributions. A specialist should still be compared with the current general model plus its normal tools on the exact task where it will be routed.

Parsing and tools: preserve more than one view of the chart

Two 2026 systems show complementary ways to make perception visible.

ExChart focuses on extracting data from charts without printed point labels. Its trained 7B model achieved 4.87 Adaptive MAPE on the authors’ 3,600-example benchmark, versus 5.94 for the much larger GLM-4.5V and 6.72 for Gemini 2.5 Flash. The authors explicitly do not treat 4.87 as automatic reliability. Their interface calibrates the axes, overlays movable points, and lets a person verify or correct the recovered values. In a 12-person study, six-point and fifteen-point charts both took about half a minute on average to correct. Participants rated final satisfaction high, but model-accuracy trust only 3.67/5. The study is small, local, and lacks a randomized direct-workflow control. Its enduring contribution is the visible correction surface: precise-looking output remains editable evidence.

ChartAgent gives GPT-4o more than 40 chart-specific tools for OCR, segmentation, color, geometry, statistics, and calculation. It runs a ReAct loop for up to 15 iterations. On ChartBench it reached 71.39 overall versus 54.53 for GPT-4o alone. Adding the system to GPT-4o improved unannotated numeric questions by 26.7 points. Chart-specific image and analysis tools beat generic image tools by 30.0 points overall.

This is unusually useful ablation evidence: it compares tool roles around the same orchestrator. It also exposes the cost of composition. In 30 inspected trajectories, tool output was correct without recovery half the time. Of the rest, 70% recovered and 30% failed—15% of all inspected trajectories. Most failures were perceptual: obstructed text, poor contrast, occlusion, segmentation, overlap, and axis problems. The paper studies single-chart question answering, not authoring, dashboards, or reader outcomes.

Making multimodal LLMs reliable chart data extractors attacks a different failure: repeated calls to the same model disagree. Across 20 TinyChart samples, 99.4% of 1,000 charts produced differing outputs and 49% produced different table dimensions. The authors align repeated tables and take per-cell medians, stopping early when results converge and using disagreement as uncertainty. This improves several open and specialist models. It costs 4–16 times more inference, and uncertainty correlates only modestly with accuracy. The method helps triage unstable cells; it does not certify them.

Current reading: the useful intermediate is plural. Keep source data, chart specification, extracted table, rendered image, localized marks, and tool observations available for different checks. A pipeline that converts the chart to one table and discards the image makes some questions easier by making others impossible.

Total-workflow cost: the pieces do not yet add up

The literature can now supply cost units, but not a specialist-versus-general total. The studies use different tasks, denominators, acceptance rules, model versions, and price schedules. Adding one paper’s model bill to another paper’s human study would create a number without a workflow behind it.

Evidence fragment Reported unit Missing from the same denominator
Self-ensembled extraction About 4–16× API inference cost; up to 20 samples per image; every reported ensembled API configuration under $15 for a full benchmark human verification, evaluator work, local compute, delivery, maintenance, readers
ExChart correction 12 people, 24 charts each; mean 31.35 seconds for six-point charts and 32.87 seconds for fifteen-point charts, with pixel-level completion enforced same-protocol current-general route, machine cost/latency, abandonment, maintenance, delivery
ChartAgent 5–7 average iterations; about 6–10 seconds for one GPT-4o call versus about 90 seconds serial or 30 seconds parallel; about $0.40 per query across 4,952 pairs human verification, evaluator labor, accepted-answer cost, maintenance, delivery, readers
ChartAgent failure audit 15% unresolved tool-level failures in 30 sampled trajectories; fallback below 10% in a separate 30-trajectory review larger denominator, human repair, consequential escapes, abandonment
VisJudge-Bench evaluation Three crowd ratings per each of 3,090 samples and three experts reviewing all samples; 11% overall and 16% sub-dimension adjustment model cost/latency, creator repair, expert time/pay, maintenance, readers
DashboardMimic evaluation 450 manually annotated interaction tasks and three experienced evaluators scoring 90 dashboards evaluator time/pay, generation and judge cost, correction, maintenance, mobile/accessibility/readers
Chart2Code-MoLA preparation and inference Reports shared-corpus training memory/time and per-chart latency for chart-specific MoE + LoRA, full-fine-tuning, and LoRA-only variants current-general route, consistent denominator and selected configuration, raw telemetry, human acceptance/correction, delivery, readers, maintenance; exact numbers quarantined after the audit below
METAL iterative chart reconstruction Same-task direct, Best-of-N, and four-role iterative results with two base models; up to five recurrences, a 512–8,192 token-budget plot, and local GPU-memory description matched calls/tokens/compute, actual early-stop and retry distribution, latency/charges/evaluator compute, human acceptance/correction, delivery, readers, maintenance
ProMCP phase profiling Adjacent six-stage token and latency attribution across three MCP deployment topologies, with initialization and tool discovery separated visualization task and specialist comparator, released task traces and historical run manifest, charges/compute, human acceptance/correction, delivery, readers, maintenance
ChartAgent tool-integrated reasoning Same-task ChartBench scheduler and component sweeps report 73.2–78.8% accuracy beside 3.0–9.8 mean tool invocations exact ablation-task count, model calls/tokens, local tool compute, reflection, latency, retries, charges, evaluator/human work, and accepted/rejected/abandoned outcomes
VisCoder2 conditional self-debug On 888 VisPlotBench tasks, VisCoder2-32B and GPT-4.1 both finish with 732 execution passes; the released round totals reconstruct to 584 versus 714 revision generations comparable tokens, accelerator/provider use, latency, charges, evaluator work, human or production acceptance, delivery, and maintenance

VisJudge-Bench permits one transparent derived range: 3,090 samples × three annotations ÷ 15 images per batch gives 618 batch-annotations. At the reported 30–60 minutes per batch and estimated $10/hour, that is approximately 309–618 hours and $3,090–$6,180 on the stated crowd-pay basis. It excludes platform fees, screening, discarded work, and all expert review. It is an evaluation fragment, not the benchmark’s total cost. Likewise, a local model’s $0 API charge is not zero compute, and ChartAgent’s cheaper GPT-4o-mini figure is a projection rather than a measured replacement run.

A defensible comparison needs one frozen task sample and acceptance rule, the current general route and specialist route under equal total budgets, and eleven visible lanes: route preparation/ownership, task/acceptance, success, inference, latency, evaluation, human correction, unresolved failures, maintenance, delivery, and reader outcomes. Keep five cost classes separate: fixed preparation, periodic ownership, marginal attempts, failure-contingent work, and downstream delivery or reader work. Declare a volume and time horizon before amortizing setup. Missing values stay null. Measured, recomputed, projected, allocated, and source-reported costs stay distinct. Report cost per eligible task, attempted task, accepted artifact, and successful reader task separately so abandonment cannot make a route look artificially cheap.

Tool calls and accuracy are not cost per accepted artifact

The new ChartAgent preprint is closer than most studies to a same-task cost-quality frontier. Its full route reports 78.6% accuracy with 6.1 mean tool invocations; removing the information-gain scheduler reports 77.9% with 8.4, removing GroupTalk 76.8% with 5.6, and removing the tool library 75.4% without a call count. The paper also varies its cost weight and cumulative budget. This supports a bounded mechanism claim: scheduler choices can move quality and one resource proxy together.

It does not price the route. Tool invocations omit the model work that selects them, local vision-model compute, final multi-expert reflection, tokens, latency, retries, evaluator work, and human correction. The exact ablation denominator is unstated, and aggregate accuracy does not identify artifacts that passed a declared acceptance contract.

Ledger v7 now holds twelve cost-relevant fragments. Zero has both equivalent observed per-arm route-wide cost and a frozen accepted-artifact denominator. It therefore defines the ratio explicitly:

reconciled route cost for the declared population and window / accepted artifact count

Rejected, abandoned, quarantined, and no-output attempts stay in the numerator. Only artifacts passing the frozen human or production contract enter the denominator. The ratio is undefined when none pass. Report eligible, attempted, candidate, accepted-without-repair, accepted-after-repair, rejected, abandoned, quarantined, no-output, delivered, and reader-successful counts beside separate per-unit ratios. Tool calls, per-attempt API averages, benchmark accuracy or F1, finish signals, and one-arm telemetry are not substitutes.

The same execution endpoint can hide different retry burden

VisCoder2, accepted at ICLR 2026, supplies a more informative retry record than a single final pass rate. Its VisPlotBench evaluation contains 888 tasks across eight visualization languages. The VisCoder2-32B route begins with 649 execution passes and GPT-4.1 with 563. After up to three conditional self-debug rounds, both finish at 732 of 888, or 82.4%.

The endpoint ties. The paths do not:

Route Initial execution pass Revision attempts by round New passes by round Total revisions Total generations
VisCoder2-32B 649 / 888 239 · 184 · 161 55 · 23 · 5 584 1,472
GPT-4.1 563 / 888 325 · 217 · 172 108 · 45 · 16 714 1,602

The paper’s aggregate tables therefore imply 130 fewer revision generations for the VisCoder2 route. They also expose a declining conditional rescue rate: 23.0%, 12.5%, and 3.1% for VisCoder2, versus 33.2%, 20.7%, and 9.3% for GPT-4.1. That makes retry burden a survival process: each round spends work on the tasks that have survived all earlier failures. A final pass rate or fixed maximum round count erases both the number of attempts and their diminishing return.

This is not a cost winner. The two routes use different models, and the paper does not report comparable task-level tokens, accelerator or provider use, latency, charges, evaluator work, or human correction. The commit-pinned implementation retains raw initial responses, but its OpenAI self-debug path prints usage and returns only response text; its local-model debug path likewise returns text without generation metadata. Judge records keep scores rather than usage or latency. Training is described as three epochs on eight H100 GPUs, without elapsed hours, energy, allocation, or labor. The released repository also does not carry the historical result bundle used for the paper tables.

Execution is only the paper’s mechanical gate. The paper itself includes cases where execution improves without corresponding visual-quality improvement. Neither route’s 732 passing programs were classified under a frozen human or production acceptance contract, delivered to intended readers, or followed through maintenance. The defensible result is therefore: same execution endpoint, different retry topology, zero accepted-cost comparison, and no route winner. A valid reopening receipt would join every task and retry to observed resources and then to accepted, rejected, abandoned, quarantined, no-output, delivered, and reader-successful states under one contract.

A matched inference budget is a real result—and still not total cost

The peer-reviewed Selective Test-Time Scaling visual-insights study is the strongest held fixed-budget near miss. It sends raw tabular data through profiling, visualization, chart checking, insight generation, pruning, and final report scoring. Across its main VIS Publication experiment, the baseline and four pruning policies use the same models and environment while total declared LLM calls stay within 2.3% of the baseline. A separate run keeps reported output tokens within 1.1%. At pruning ratio 0.6, the mean proxy-judge score rises from 61.64 to 65.86 while the number of generated final reports falls from 1,435 to 578.

That supports a narrow mechanism result: under one pipeline and declared partial inference budget, early pruning can allocate search more effectively. It does not compare a specialist route with a direct current-general route, and “final report” means a generated candidate rather than work accepted under a human or production contract.

The official implementation also keeps the cost boundary visible. Stage counters are assumptions rather than reconciled provider records; plot execution can call a code rectifier without adding that call to the budget; metadata failures return zero budget; errored downstream workers are excluded from aggregate totals; and token records contain completion tokens only. The repository describes local result manifests but does not ship the experiment traces at the audited commit. These are release-accounting limits, not evidence that the private experiment failed.

Four expert annotators calibrate the proxy judge on score-stratified reports from two datasets. That is stronger than an unvalidated automated rating, but it still does not classify every task as accepted, rejected, abandoned, quarantined, or no-output. DV-World adds a separate same-harness interaction- charge versus capability view. Neither fragment supplies the full numerator or denominator.

The operational distinction is now three receipts, not one:

Receipt What it can establish What it cannot establish alone
Matched declared partial budget Two policies were allotted closely matched calls, output tokens, or another named slice. Actual all-resource equality, failure cost, people, or delivery.
Observed route-wide use Every eligible task retains actual machine, evaluator, human, failure, preparation, delivery, and maintenance events. How many outputs passed the real acceptance bar.
Cost per accepted artifact Reconciled route-wide cost is divided by artifacts passing one frozen contract. Reader success or later maintainability unless those are inside the contract and window.

Practitioners should keep repair and discarded candidates beside the task. BI leaders should ask vendors for all three receipts. Researchers should release task traces and preserve failed branches. Newsrooms should define editorial, source, accessibility, and reader gates before timing. Product teams should instrument request and retry identity rather than increment assumed stage counters. Educators and accessibility specialists should put representative use inside acceptance. Sponsors should read the shortest defensible answer as: one useful fixed-budget allocation result, zero of twelve total-cost-per- accepted-artifact comparisons, and no route winner.

Equal rounds are not an equal budget

METAL is useful same-task mechanism evidence. At five recurrences, its authors report average automatic F1 of 51.78% versus 40.45% for direct prompting and 43.13% for Best-of-N with Llama 3.2-11B; with GPT-4o, the corresponding values are 86.46%, 81.26%, and 82.32%. The system separates generation, visual critique, code critique, revision, deterministic verification, and automatic evaluation. The authors also say explicitly that it costs more than direct prompting.

The matched n = 5 rows are not an equal-resource comparison. In the official implementation at commit 41832c1, Best-of-N performs five generation calls. The METAL path performs one initial generation plus visual-critique, text-critique, and revision calls in each recurrence—up to 16 model calls at five recurrences, plus a possible retry. The paper describes four calls per iteration, so a historical run receipt is needed to reconcile the paper and code definitions. The public wrapper and logs do not retain per-arm provider request IDs, input/output token use, call timestamps, latency, charges, accelerator-seconds, energy, or evaluator compute. The public runners also recommend one process for METAL and eight for baselines, preventing elapsed time from being compared without observed concurrency and occupancy.

This does not invalidate METAL’s automatic performance result. It means the result cannot establish equal-budget efficacy or total cost. The corrected comparison has two resource records:

Receipt Required content Decision use
Planned cap Calls/retries, tokens and image units, tool/evaluator work, accelerator identity/time, latency and concurrency, charges, human minutes, and fixed-to-downstream allocations Freeze the resource rule before quality is known.
Observed use Route/task/attempt/event IDs, timestamps, stop reason, actual use in every declared dimension, value status, and source receipt Distinguish equal resources, unequal-but-costed routes, missing use, and conflicting use.

Matching iterations, candidates, provider rates, nominal token caps, or model labels alone satisfies neither receipt. A provider rate without observed usage is a price schedule, not workflow cost.

A machine total still hides the bottleneck

ProMCP is adjacent evidence about tool-mediated agents, not a chart-specific cost comparison. Across 155 MCP-Bench and MCP-Universe tasks, its authors attribute tokens and latency to user context, planning, tool calls, tool results, context updates, and final synthesis; they record session initialization and tool discovery separately. Customized-client configurations concentrate much of their measured burden in planning and schema injection, while the off-the-shelf Claude Desktop configuration concentrates most of its MCP-Bench latency in final-answer synthesis. The point is structural: changing the host, client, model placement, context policy, or streaming behavior can move the expensive phase even when the visible workflow still looks like “model plus tools.”

The limits are equally important. The off-the-shelf path is reconstructed from exported conversation timestamps and cannot expose hidden retries. The study uses one workstation and lightweight-to-moderate tools, so database, remote data, rendering, OCR, or local vision work can produce a different bottleneck. It reports tokens and latency rather than accepted visualization work or total human and machine cost.

The official release at commit 524b4ad provides useful event fields but no task-level traces or historical effective-settings manifest from which the paper tables can be recomputed. Its advertised benchmark entry point imports a benchmark.runner module that is not present. Checked-in defaults also differ from the paper settings on rounds, caching, and concurrency. These findings do not disprove the paper results; they prevent the public repository from being treated as a reproduced run receipt.

Ledger v7 therefore retains one topology-bound event record beneath the existing workflow lanes. It retains cold or warm state, schema-set hash, host/client/ server/model/tool identity, transport, cache, streaming, concurrency, timeout and retry policy; stage timestamps, tokens and payload footprint; provider or accelerator receipt; stop reason, outcome, value status, and source receipt. Hidden phases remain missing rather than being inferred from user-visible messages. Phase telemetry makes the machine portion inspectable. It does not replace human correction, accepted delivery, accessible use, reader outcome, or maintenance evidence.

A cost table still needs an integrity gate

Chart2Code-MoLA is useful because it separates training resources from per-chart inference on one synthetic chart-to-code setup. It is not a current- general comparison, and a direct audit found four conflicts in the paper’s own record. Its five displayed split rows sum to 110,000 training, 17,000 validation, and 18,000 test examples, while the total row says 112,000/24,000/24,000. A 0.45-versus-0.40-second latency difference is 12.5%, not the stated 5%. The peak-memory table prints an unexplained 4.84 delta for 12.3 versus 15.0 GB. The implementation section specifies top-2 routing, attention + MLP, and alpha 16, while the ablation/conclusion names probabilistic routing, attention + output, and alpha 32.

Those conflicts are present in the original PDF, not introduced by text extraction. They do not show that the method fails. They do mean its exact success and efficiency numbers cannot enter a route ranking until a corrected record or reproducible receipt binds one data denominator and configuration. The ledger now checks route identity, denominator arithmetic, percentages, configuration binding, and value status before admitting telemetry. Passing that preflight makes a receipt usable; it does not establish efficacy or a winner.

For practitioners, time the path from route preparation through accepted delivery and keep cold start, discovery, planning, tool/result transfer, context update, and synthesis separate. For BI and analytics leaders, require a topology card and phase receipt beside the total so the pilot names what is local, cloud, cached, streamed, concurrent, or hidden. For editors, source checking, correction, mobile and keyboard proof, and publication belong in the ledger. For researchers, release task traces, effective run manifests, observability state, denominator recomputations, and a later-change receipt. For product and engineering teams, instrument retry edges, effective settings, usage and charges before optimizing. For learners and accessibility specialists, keep assistive-path preparation and reader use separate from creator speed. For sponsors, the short answer is still that no cost winner is known.

The comprehensive audit and machine-readable ledger remain in private research custody and make that comparison answerable. They are a protocol, not a completed comparison: no route ran, no budget was authorized, and no cost winner exists.

Quality and integrity critics: direct progress, narrow authority

VisJudge is the closest thing in the current literature to a visualization-specific visual critic. Its benchmark contains 3,090 visualizations: 1,041 single charts, 1,024 multiple-chart compositions, and 1,025 dashboards. Three crowd annotators rated each sample; three visualization experts reviewed all annotations. The rubric adapts six dimensions—fidelity, semantic readability, insight discovery, design style, visual composition, and color harmony—to the chart.

The resulting 7B model, trained from Qwen 2.5-VL with supervised and reinforcement learning, reached .421 mean absolute error and .687 correlation on the held-out test. GPT-5 reached .553/.428, GPT-4o .610/.482, and Claude 4 Sonnet .622/.465. This is a meaningful specialist win on a directly relevant task.

It is not a final acceptance gate:

Misviz measures a narrower but important kind of critique: detecting 12 misleading-design categories. It includes 2,604 real-world visualizations and 57,665 synthetic charts. On real charts, GPT-o3 led at 80.0 F1 and 58.8 exact match; the image-plus-extracted-axis classifier reached 60.9 F1 and 12.3 exact match. On controlled synthetic charts, the rule-based and classifier systems became competitive. A linter with ground truth axes achieved 99.7 precision but only 52.2 recall. The real-world result favors broad general reasoning; the synthetic result favors encoded rules. Neither captures narrative deception, domain knowledge, or the full published taxonomy of misleading techniques.

VisDeception uses 1,600 paired charts: the same data rendered faithfully and with one of eight deceptive tactics. A multi-stage system first extracts structured metadata and then reasons about the chart. It reduced Gemini 2.5 Pro’s paired error from .099 to .024 and GPT-4o’s from .502 to .255. It worsened Gemini 2.5 Flash from .109 to .135 and Qwen3-VL-2B from .251 to .275. The generated paired design isolates causality well; it does not represent messy surrounding prose, captions, or real-world author intent.

Current reading: use a specialist critic as one bounded witness. Pair a perceptual-quality rubric with deterministic data/spec tests, an integrity taxonomy, accessibility checks, and contextual review. Its output should be a diagnosis with evidence, not a single universal score.

Grounding: the critic must point to what it means

A critique becomes more useful when it can locate the mark, label, axis, or legend that caused the judgment. ChartREG++ tests this directly with 850 chart images, 3,400 referring expressions, 18 element types, and point, box, and mask outputs. The references include textual, data, and visual cues.

Strong general multimodal models remained below 50 F1 for point and box grounding. For masks, a zero-shot SAM 3 candidate pipeline plus language alignment reached 25.01 and 21.84 F1 on the two evaluation sources. Synthetic chart-specific Mask2Former candidates plus the same alignment idea reached 52.55 and 49.85. This is evidence that generic open-vocabulary segmentation does not automatically resolve fine, overlapping chart marks. It is not a fair zero-shot comparison: the chart candidate model was trained specifically for the domain.

The chart-specific masks transferred to a modified and verified real-chart set, especially for lines, but the benchmark remains static and concentrated on common Matplotlib-style elements. Candidate quality and language alignment are still separate bottlenecks.

Current reading: grounding is a strong candidate for the critique stack because it can connect a finding to visible evidence and support a repair check. The missing experiment is not another mask score. It is whether grounding helps a system identify a real defect, edit the right object, and verify that the edit fixed the defect without changing something else.

OCR and document models: useful sensors, not quality judges

Three recent releases matter because many real visualizations arrive inside reports, dashboards, slide decks, and screenshots rather than as clean chart images.

These systems may recover labels, reading order, chart regions, or a structured graphic before a critic reasons. That is a plausible architecture role, not a demonstrated outcome. None of these studies compares the same visualization critic with and without the parser, measures whether source-data errors are found, or tests comprehension after the parsed evidence is used.

Current reading: add a document specialist when OCR or layout is a measured bottleneck. Evaluate the entire downstream task. Do not convert a document benchmark lead into authority over chart quality.

General-model capability is improving, but not uniformly

Two human-comparison studies help locate the moving baseline.

CHART-6 adapted six independently developed human visualization-literacy tests to eight vision-language models. The strongest 2024-era model, GPT-4V, remained below human performance on several tests, and model error patterns were far from the human noise ceiling. The work also found that invalid answer formatting materially changed scores—a harness issue that can masquerade as visual reasoning.

A July 2026 scientific visualization literacy study tested six newer models on 49 items spanning 18 scientific visualizations and illustrations, with 485 human nonexperts. Gemini 3.1 Pro Preview reached 88.6% overall; Claude Opus 4.6 reached 75.3, GPT-5.4 75.1, and the human table reports 75.6. The paper’s narrative gives a slightly different 75.9 human figure. The result does not mean Gemini has universal scientific visualization literacy: some task categories contain only one or two items, and fine quantitative reading, flow direction, and unsupported encodings remain difficult.

The baseline has plainly moved. What has not disappeared is the need for specialized evidence, executable checks, and task-specific evaluation. Better general models make generic “chart expertise” packages less defensible; they make narrow specialists with inspectable added information more valuable.

Chartography makes the remaining professional reading gap concrete. Its 100 deliberately difficult tasks were authored by practitioners across 12 domain labels and independently verified by three experts each. The best of 30 frontier configurations reached 45.0% mean pass@1. More reasoning helped most paired configurations by a median 4.5 percentage points, but often extended an initial misreading of sparse axes, 3D projection, contours, or domain conventions. Because the benchmark screened for difficulty, the result is not a failure rate over ordinary charts. It is evidence that professional environment and visual form remain routing variables.

Stateful dashboards: a screenshot is not the whole artifact

Dashboard2Code adds 180 Plotly Dash dashboard-code pairs, 20 visualization types, eight callback patterns, and 450 interaction tasks. The best reported configuration scored 79.4 overall and 64.2 on the hardest interaction level. DOM access raised one model’s successful UI exploration from 20.61% to 64.58%. Ninety generated dashboards were also scored by three visualization-experienced graduate evaluators; the final automatic metric correlated .781 with those human ratings.

The benchmark exposes failures a static critic misses: hidden state, cross-control dependencies, and factually wrong transformations can coexist with a visually responsive interface. It also draws a firm scope boundary: Plotly Dash only, fixed 1920×1080 viewport, no animation or popups, and no mobile, keyboard, assistive-technology, or reader evaluation.

Context changes what “better” means

A journalism graphic, a BI dashboard, a scientific figure, and an accessible description do not share one quality function.

The common spine is the same—intent, source, representation, rendering, verification, delivery, reader—but the specialist and acceptance evidence must be delegated by context.

Where the literature is converging

Preserve structure and pixels

Tables, chart specifications, OCR blocks, masks, and images are complementary. The recurring architecture exposes intermediate evidence instead of hiding all perception and reasoning inside one completion.

Separate perception from judgment

The papers increasingly isolate axis recovery, mark localization, calculation, deception detection, and perceptual rating. This makes failures diagnosable and allows a system to assign the smallest relevant specialist.

Use executable operations for calculations and checks

Program-of-thought, chart analysis tools, linters, and deterministic source/spec validation reduce reliance on visual arithmetic. They remain bounded by their inputs and taxonomies.

Keep correction mixed-initiative

ExChart’s overlay, ChartAgent’s visible tools, and ensemble disagreement all make uncertainty inspectable. They support a human deciding whether and how to repair rather than merely accepting a confident answer.

Test on harder distributions

ChartQAPro, CharXiv, Misviz’s real split, ChartREG++, scientific visualization literacy tests, and accessibility-context ablations all move beyond clean synthetic single charts. The hardest findings are often transfer failures, not another point of in-distribution progress.

Where approaches legitimately diverge

What has performed poorly or caused harm

These are not theoretical warnings; each has direct evidence in the reviewed work.

Technique or assumption Negative finding Practical consequence
Select a model because it led an older chart benchmark ChartGemma and TinyChart collapse on ChartQAPro; DePlot and TinyChart trail frontier general models on WB-ChartExtract Re-run same-task baselines after model and benchmark changes
Treat a table as a lossless chart representation DePlot loses color, orientation, and visual attributes and trails MatCha on PlotQA Preserve the rendered visual and explicit encodings alongside extracted values
Train on synthetic charts and assume real-world transfer Misviz method rankings change sharply between synthetic and real charts Maintain separate synthetic diagnostic and real acceptance sets
Add generic visual feedback Generic Qwen feedback often lowers downstream chart quality in the VisJudge experiment Evaluate the critic itself and expose its rubric
Add a vision caption by default In three small product-owned Vizier comparisons, the captioned arm won 0/3, 1/5, and 1/5 cases while no-caption won 1, 3, and 4 Treat captions as routed, fallible sensor output; retain them only when they improve a named downstream decision
Apply one mitigation pipeline to every model VisDeception’s structured mitigation improves several models and worsens Gemini 2.5 Flash and Qwen3-VL-2B Route and ablate by model/version; preserve a no-mitigation baseline
Use a general segmenter for fine chart marks SAM 3 candidate grounding remains weak on ChartREG++ Use chart-aware candidates when grounding is a measured bottleneck
Trust a single extraction 99.4% of TinyChart cases varied across 20 samples; 49% changed table dimensions Surface instability and require correction for consequential use
Treat ensemble disagreement as confidence Correlation with accuracy is only about −.30 to −.37 and inference costs grow 4–16× Use it as escalation evidence, not a probability of correctness
Treat an automatic judge or preference as reader comprehension VisJudge revisions are automatically evaluated; the BLV tactile + text + LLM condition was strongly preferred but did not improve measured accuracy Add human expert and reader acceptance, and keep preference, mental model, comprehension, and decisions separate
Treat a fixed desktop dashboard as responsive evidence Dashboard2Code measures stateful Plotly Dash behavior at 1920×1080 and excludes animation and popups Test multiple viewports, responsive reflow, animation, popups, keyboard paths, and assistive technology

A reference architecture for critique

The research supports a layered system with explicit evidence ownership:

Layer Primary evidence Appropriate mechanism Authority boundary
Intent and context audience, task, claim, stakes, environment, device human brief + general reasoner determines which quality dimensions matter
Source truth data, units, joins, transformations, semantic definitions deterministic queries, tests, provenance authoritative for values; no image model can replace it
Chart structure specification, scales, encodings, labels, states schema checks, linter, chart parser validates declared structure, not reader interpretation
Rendered evidence pixels, text, geometry, marks, occlusion OCR, chart extraction, grounding/segmentation identifies visible evidence and discrepancies
Task-specific critique integrity taxonomy, perceptual rubric, domain conventions specialist critic + deterministic rules limited to tested dimensions and distributions
Broad synthesis surrounding prose, purpose, exceptions, tradeoffs current general multimodal model integrates context; should cite upstream evidence
Delivery responsive states, interaction, keyboard, screen reader, latency browser/device replay and accessibility tests validates the artifact people receive
Acceptance comprehension, decision quality, harm, editorial or domain judgment intended readers and accountable humans final authority for consequential communication

This architecture does not require every layer for every chart. A simple local plot may need source checks and a rendered inspection. A public dashboard may need all eight. The important design choice is that no component silently gains authority over a layer it cannot observe.

Adoption signals: evidence of attention, not effectiveness

Several projects have public repositories. As captured on 2026-08-14:

Project GitHub stars Forks What it indicates
SAM 3 11,315 1,704 substantial attention to a broad segmentation model, not chart-grounding quality
dots.ocr 9,069 802 strong interest in an open OCR/structured-document model
HunyuanOCR 1,920 149 material attention to a lightweight OCR family
VisJudgeBench 124 6 early research uptake for a new visualization critic
ChartQAPro 45 7 a benchmark repository, where stars are especially weak as use evidence

Stars and forks can prioritize integration research. They do not establish downloads, active deployments, retention, task fit, reproducibility, or output quality. SAM 3’s broad adoption signal is particularly easy to misread: ChartREG++ directly finds that a generic SAM 3 candidate pipeline is not enough for fine chart-element grounding.

Gap ledger

Question Status Evidence in hand Evidence still needed
Do chart specialists beat strong general models on realistic OOD tasks? Answered: no class-wide ordering ChartQAPro and WB-ChartExtract show severe specialist transfer gaps; narrow author-benchmark wins also exist Reopen after independently replicated dominance across realistic tasks
Do chart-specific tools add value? Answered for chart QA ChartAgent’s same-orchestrator ablations show large gains over no tools and generic tools Replication on newer models, dashboards, generation, and critique
Can a quality critic improve final charts for people? Partly answered VisJudge improves expert-rating prediction and automatic downstream scores Human expert and reader evaluation of the revisions
Does stronger OCR/document parsing improve critique? Open PaddleOCR-VL, HunyuanOCR, and dots.mocr improve upstream recovery End-to-end ablation with the same critic and acceptance task
Does chart grounding improve defect localization and repair verification? Partly answered ChartREG++ shows chart-aware masks improve localization Diverse real charts and a repair-verification outcome
Can model disagreement drive safe escalation? Partly answered Repeated sampling improves extraction and reveals instability Calibrated thresholds across models, distributions, costs, and stakes
Do specialist-assisted workflows improve reader or accessibility outcomes? Partly answered A 12-participant BLV study finds strong preference and spatial-model benefit for tactile + text + LLM, but no accuracy lift and no isolated chart-vision specialist Delivered-chart specialist ablation on comprehension, decisions, mobile use, and assistive technology
What happens on interaction, animation, responsive states, and dashboards? Partly answered Dashboard2Code measures callbacks and hidden state in 180 fixed-desktop Plotly Dash applications; VisJudge dashboard correlation is weaker Multi-viewport stateful benchmark with browser traces, responsive rendering, animation, keyboard paths, and assistive technology
Which specialist belongs in which environment? Partly answered The role typology explains why routing should differ; Chartography shows professional domain conventions and visual forms materially affect performance Common journalism, BI, science, operations, and accessibility routing comparison
What is the total human and machine cost? Partly answered: ledger corrected, comparison not run Twelve fragments expose partial preparation, inference, latency, correction, evaluator, failure, call-topology, phase, interaction-charge, retry-survival, and tool-use units; one matches a declared partial inference budget, while zero has both equivalent observed per-arm route cost and a frozen accepted-artifact denominator; ledger v7 keeps the three receipt levels separate Run current-general and specialist routes on the same tasks with one acceptance rule, equivalent observed usage, all outcome states, preparation and amortization, delivery, reader task, and later maintenance event
Can the held studies support cost per accepted artifact? Answered: zero of twelve held comparisons Selective TTS is the strongest fixed-budget near miss; VisCoder2 adds a reconstructable same-endpoint retry topology, but neither supplies equivalent observed route use plus task-level acceptance Reopen after a held study or authorized run reports equivalent arm receipts and accepted, rejected, abandoned, quarantined, and no-output states under one contract

This ledger is the practical boundary of the review. The first two questions have bounded answers. Six have meaningful component or adjacent evidence but remain open at the workflow or reader level. Two remain direct corpus gaps.

1. Establish the moving baseline

Use the current general multimodal model with its normal harness. Measure the actual local task, include hard and unanswerable cases, and record value accuracy, defect detection, latency, cost, and human correction. Re-run after a material model or harness release.

2. Add deterministic source and specification checks

Before another model, verify values, units, joins, transformations, scales, labels, and expected states from authoritative inputs. This establishes which errors actually require vision.

3. Add one perception specialist at a measured bottleneck

If OCR fails, test a document parser. If values cannot be recovered, test a chart extractor. If a critique cannot point to the right mark, test chart grounding. Keep the rest of the system fixed and score the end-to-end task, not the component’s published benchmark.

4. Compare critic roles separately

Run a perceptual-quality critic, an integrity taxonomy, deterministic rules, and the general model on the same charts. Measure agreement with accountable experts by dimension. Do not collapse the outputs into one score until each component’s false-positive and false-negative costs are known.

5. Test repair, not only detection

Require the system to localize a defect, propose or perform the smallest edit, and prove that the targeted defect improved without regression. Compare a grounded critic with a text-only critic and preserve the original artifact.

6. Evaluate ensemble routing under a fixed budget

Compare single pass, repeated sampling, and escalation to a stronger model or human. Set one total inference, latency, and reviewer-time budget. Evaluate calibration and abandoned work, not only average extraction gain.

7. Finish at the delivered reader experience

Test desktop and phone layouts, interaction states, keyboard access, screen reader output where relevant, comprehension, confidence, and decision quality. A better screenshot score is an intermediate result, not the goal of data visualization.

Method and limits

This is a purposive deep reading of 26 papers plus primary model cards, repository records, and three product-owned historical caption-comparison receipts. The local receipts are self-reported, small, and kept separate from independent research. The review emphasizes recent 2025–2026 work while retaining the 2023–2024 systems needed to explain the technique history. Papers were read in full where readable text was captured. Numerical comparisons appear only when reported on the same test and metric. Author-benchmark results are labeled; automatic judges, small human samples, preprint status, and static/synthetic scope are carried into the interpretation.

The review is not exhaustive. Most evidence is English-language, static, and same-team. Frontier models move faster than peer review and are not evaluated on every specialist benchmark. Repository popularity is not usage. No local deployment, cost, privacy, licensing, or hardware evaluation was performed.

The central conclusion is therefore intentionally conditional: route a specialist when it supplies a tested piece of evidence that the baseline lacks, and keep its authority no broader than that evidence.

Update log

Read or download the Markdown source