Specialized vision models for data visualization: what they can—and cannot—verify
Status: research snapshot, evidence cut 2026-08-16. Recheck by 2026-11-16, or earlier after a major general-model release, independent replication, or specialist-assisted reader study.
Companion: The state of AI-assisted data visualization research.
This report examines specialized vision models and systems that might help an AI understand or critique a data visualization. It covers chart question answering, chart-to-table extraction, document OCR, chart-element grounding, misleading-chart detection, perceptual quality judgment, and verifier or repair loops. It asks where specialization adds something beyond a strong general multimodal model—and where the available evidence does not support that claim.
Executive assessment
There is no specialist model that can serve as a trustworthy final judge of a data visualization. There are useful specialists for different parts of the problem.
A chart parser can recover values without knowing whether the chart supports its headline. A grounding model can locate a line or legend without verifying the source data. A misleading-chart detector can identify a truncated axis while missing a false annotation. A learned quality critic can reproduce expert ratings of composition and readability without knowing whether readers understand the result. These are different jobs with different evidence.
The findings support six conclusions:
- “Specialized” is not a quality ranking. Chart-specific models that led older benchmarks performed poorly when ChartQAPro introduced realistic, harder, and out-of-distribution charts. On a separate 2026 extraction benchmark, a frontier general model substantially outscored DePlot and TinyChart.
- Specialized tools can still add large gains to a general reasoner. In ChartAgent, a chart-specific tool system improved the same GPT-4o model by 26.7 percentage points on unannotated numeric questions; chart-specific tools beat generic image tools by 30.0 points overall.
- The most direct specialist critic is promising but not a final quality gate. VisJudge predicts its expert-adjudicated visualization ratings better than the tested general models. Its chart-revision gains were scored by another model, not by readers or human experts, and it cannot verify source fidelity from an image alone.
- Better perception is useful evidence, not a complete critique. New OCR, structured-graphics, and grounding models can recover text, layout, values, and marks. None of the captured evaluations shows that a stronger parser by itself produces better end-to-end judgments about communication, integrity, or reader outcomes.
- The defensible architecture is composed and context-dependent. Start from source data and chart specification, run deterministic checks, add the smallest specialist that addresses a measured perception failure, retain a current general model for broad reasoning, and preserve human acceptance. When comprehension, decisions, mobile use, or accessibility matter, test the delivered experience with the intended readers.
- Machine totals need phase and topology. Adjacent ProMCP evidence shows measured burden moving between planning/schema injection and answer synthesis as client topology changes. Preserve cold start, discovery, planning, tool/result transfer, context update, and synthesis separately; this is a measurement rule, not a visualization cost winner.
The specialized-vision gap in the broader field review should therefore move from open to partly answered. The field now has credible evidence for several component roles and clear negative findings. It still lacks a complete routing study, independent replications, end-to-end cost evidence, and human reader outcomes that isolate the specialist’s contribution.
“Vision for charts” contains six different problems
| Role | Input and output | What success means | What it cannot establish alone |
|---|---|---|---|
| Chart QA and reasoning | chart image + question → answer | correct answer on a declared chart distribution | source fidelity, overall quality, reader comprehension |
| Chart parsing | chart image → table, values, or structured graphic | recovered labels, series, geometry, and numbers | whether the visual encoding or message is appropriate |
| Document OCR and layout | page or screenshot → text, regions, tables, formulas, reading order | accurate document structure under realistic degradation | chart semantics, insight, integrity, or aesthetic quality |
| Grounding and segmentation | image + referring phrase → point, box, or mask | the requested mark or region is localized | whether the mark is true, legible, or meaningful |
| Integrity and quality critique | chart image, sometimes with context → defect labels or ratings | agreement with a defined taxonomy or human rating protocol | dimensions excluded from the rubric or unavailable source evidence |
| Verifier and repair component | repeated outputs, intermediate evidence, or candidate revision → score, uncertainty, or edit | errors become detectable and repair improves a held-out outcome | correctness outside the tested loop or after delivery |
This typology is the organizing principle for the rest of the report. Model size, release date, and the word “chart” in a benchmark name are secondary to the role being measured.
A short history: progress and harder tests arrived together
The apparent story changes depending on whether the timeline tracks model scores or the questions researchers learned to ask.
| Date | Work | What changed | What later evidence qualified |
|---|---|---|---|
| Dec. 2022–2023 | DePlot | Converted a chart image to a linearized table so a language model could reason over values; DePlot + LLM reached 67.6 on ChartQA’s human split versus 38.2 for the prior MatCha system. | Tables discard color, orientation, geometry, and other visual evidence; DePlot scored 23.06 RMSF1 on a harder 2026 extraction test. |
| Jan.–Apr. 2024 | ChartAssistant, TinyChart | Added staged chart-table alignment, high-resolution visual token merging, and program-of-thought arithmetic in compact models. TinyChart’s program execution raised calculative accuracy from 56.64 to 78.98 on its older evaluation. | Both families transferred poorly to ChartQAPro; older benchmark leadership is not stable evidence of realistic chart understanding. |
| Jun.–Sep. 2024 | CharXiv, ChartGemma, EvoChart | Scientific charts became a distinct test; synthetic instructions and self-training expanded chart-specialist training data. | Generated training and author benchmarks cover bounded chart distributions. ChartGemma scored 9.80 with chain-of-thought on ChartQAPro. |
| Apr.–May 2025 | ChartQAPro, CHART-6 | Replaced templated questions with human-written, realistic, conversational, hypothetical, fact-checking, and unanswerable tasks; compared model errors with six human visualization-literacy tests. | Best reported ChartQAPro result was 55.81 versus an 85.02 human baseline; no single benchmark represents all visualization literacy. |
| Aug.–Oct. 2025 | Misviz, VisJudge | Shifted from answering questions to detecting misleading designs and predicting expert-adjudicated quality across single charts, multiple charts, and dashboards. | Real and synthetic misleading-chart rankings differ. Quality prediction still lacks raw-data fidelity and reader-outcome evidence. |
| Jan.–Mar. 2026 | PaddleOCR-VL-1.5, dots.mocr | Sub-billion to 3B models reported strong document parsing; structured graphic and SVG recovery became an explicit output. | Document parsing accuracy is an upstream capability, not a visualization-quality score. Some results rely on author benchmarks or model judges. |
| Apr.–May 2026 | ExChart, ChartAgent, ChartREG++, self-ensembling extraction | Systems added chart-calibrated human correction, specialist tools, mark-level grounding, repeated sampling, and uncertainty. | Tools fail, candidate masks constrain results, inference grows 4–16×, and uncertainty remains only modestly correlated with correctness. |
| Jul.–Aug. 2026 | HunyuanOCR-1.5, VisDeception, Dashboard2Code, Touching or Chatting, scientific visualization literacy, accessible chart descriptions, CURV, Chartography | Lightweight document parsing improved; deception mitigation, stateful dashboard reconstruction, BLV learning, professional chart reading, newer frontier-model literacy, context ablations, and explicit grounded-reasoning curricula arrived. | Multi-stage mitigation can hurt some models. Stateful evaluation remains fixed-desktop; the BLV study finds preference and spatial-model benefit without an accuracy lift; several works are recent preprints. |
The field has improved over 36 months, but the timeline does not support one smooth capability curve. As models improved, benchmarks expanded from clean single charts and short answers to scientific figures, dashboards, unanswerable questions, deceptive designs, element localization, and accessible descriptions. A score from 2023 and a score from 2026 often answer different questions.
What the strongest same-test comparisons show
The table below deliberately avoids cross-benchmark ranking. Every comparison within a row uses the benchmark, metric, and condition reported by the same study.
| Problem and benchmark | Specialist or intervention | General or base comparison | Supported reading |
|---|---|---|---|
| Realistic chart QA — ChartQAPro, chain-of-thought accuracy | ChartGemma 9.80; TinyChart 14.24 | Claude Sonnet 3.5 55.81 | Historical chart specialists did not transfer to this harder distribution. |
| OOD chart-to-table — WB-ChartExtract, single-pass RMSF1 | DePlot 23.06; TinyChart 28.61 | Gemini 2.5 Pro 87.83; Claude Opus 4.6 60.99; GPT-5.1 51.26 | General frontier models can dominate specialists on a new extraction distribution. |
| Chart QA with tools — ChartBench overall accuracy | ChartAgent 71.39 | GPT-4o 54.53; Phi-3-Vision 55.32 | A general orchestrator with chart-specific tools can outperform standalone systems. |
| Chart QA tool ablation — ChartAgent overall accuracy | chart-specific image + analysis tools: +30.0 points over generic image tools | same orchestrator and benchmark | The gain comes from task-relevant executable tools, not simply adding an agent loop. |
| Quality rating — VisJudge held-out MAE / correlation | VisJudge-7B .421 / .687 | GPT-5 .553 / .428; GPT-4o .610 / .482; Claude 4 Sonnet .622 / .465 | A small task-specific critic better matches this expert-adjudicated rubric. |
| Misleading-chart detection — Misviz real-world F1 | image + extracted-axis classifier 60.9; linter + extracted axis 36.1 | GPT-o3 80.0; GPT-4.1 79.6 | General MLLMs lead on heterogeneous real-world visualizations. |
| Misleading-chart detection — Misviz synthetic F1 | linter + predicted axis 67.9; image + axis classifier 65.2 | GPT-o3 66.9 | Rules and classifiers remain competitive on the controlled distribution they encode. |
| Chart grounding — ChartREG++ point/box F1 | best reported point F1 remained below 50; chart-specific mask candidates + alignment reached 52.55/49.85 mask F1 on two sets | SAM 3 zero-shot candidates + alignment reached 25.01/21.84 | Fine chart marks remain hard; chart-aware candidates help, but this is not a fair zero-shot model contest. |
| Deception mitigation — VisDeception error rate | structured-metadata multi-stage system reduced Gemini 2.5 Pro .099→.024 and GPT-4o .502→.255 | same models without mitigation | Explicit decomposition can help substantially, but not universally. Gemini 2.5 Flash worsened .109→.135. |
| Explicit grounded-reasoning curriculum — public benchmarks | CURV-7B improved its Qwen base: ChartQAPro 29.77→32.80 and CharXiv 32.50→36.70 | same base checkpoint | Training explicit grounding improves the selected base, but the absolute hard-benchmark gap remains and the preprint lacks external replication. |
These results settle one question and open another. There is no class-wide specialist advantage. There are strong component-level interventions. The system-design problem is to route each only where its information or executable capability adds to a current baseline.
Chart reasoning: compact specialists improved the old tasks, not the whole problem
DePlot established a durable idea: do not ask one model to perceive the chart and calculate the answer in a single hidden step. It first converts the image to a table, then lets a language model answer over that table. The decomposition made values inspectable and produced a large ChartQA gain. It also created an information bottleneck. A table preserves values and labels; it does not preserve color, line style, spatial emphasis, annotations, or interactive state.
ChartAssistant and TinyChart developed the same general direction differently. ChartAssistant first aligned charts with tables, then performed multi-task instruction tuning. TinyChart compressed high-resolution visual tokens and emitted Python for arithmetic questions. Both show that explicit alignment and executable calculation can help. They also show why an older benchmark score should not be mistaken for a lasting model choice.
ChartGemma used a 3B PaliGemma model and 122,857 charts, with Gemini-generated instructions. An audit of 100 generated instructions found 82% accurate and 8% partially accurate; only two volunteers performed the human evaluation. It reported strong results on older benchmarks, then scored 9.80 with chain-of-thought on ChartQAPro. This is not evidence that ChartGemma was poorly designed. It is evidence that benchmark fit was much narrower than the label “chart understanding.”
EvoChart scaled synthetic self-training to 1.6 million question-answer pairs. Its 4B model reached 54.2 on the authors’ EvoChart-QA test versus 49.8 for GPT-4o. That is a legitimate same-test result, with two cautions: the paper gives inconsistent 625/650 counts for the benchmark charts, and the test covers four common chart types and basic comprehension. On modified familiar charts, GPT-4o fell from 85.7 to 45.8 and TinyChart from 83.6 to 45.8. Small visual changes exposed brittle pattern matching in both general and specialist systems.
ChartQAPro is important because it changes the task. Its 1,341 charts come from 157 sources and include dashboards, infographics, and multiple charts. Its 1,948 human-written and verified questions include conversations, hypotheticals, fact checking, and unanswerable cases. The best reported result, Claude Sonnet 3.5 with chain-of-thought, was 55.81 versus an 85.02 human baseline. The human estimate comes from one expert graduate student on 50 questions per category, so it is not a population norm. The benchmark is still static and excludes interactive dashboards. It is a better transfer test, not the final definition of chart literacy.
The August 2026 CURV preprint trains explicit visual grounding before harder reasoning. CURV-7B improves its Qwen 2.5-VL-7B base across ChartQA, ChartQAPro, CharXiv, ChartMuseum, MathVista, and MMMU-Pro. That breadth makes it worth watching. It remains a new, unreplicated preprint; some answer evaluation uses GPT-4.1-mini, and level-three multi-chart performance remains at or below 42.54.
Current reading: chart-tuned training is most credible when it improves a current base across several held-out distributions. A specialist should still be compared with the current general model plus its normal tools on the exact task where it will be routed.
Parsing and tools: preserve more than one view of the chart
Two 2026 systems show complementary ways to make perception visible.
ExChart focuses on extracting data from charts without printed point labels. Its trained 7B model achieved 4.87 Adaptive MAPE on the authors’ 3,600-example benchmark, versus 5.94 for the much larger GLM-4.5V and 6.72 for Gemini 2.5 Flash. The authors explicitly do not treat 4.87 as automatic reliability. Their interface calibrates the axes, overlays movable points, and lets a person verify or correct the recovered values. In a 12-person study, six-point and fifteen-point charts both took about half a minute on average to correct. Participants rated final satisfaction high, but model-accuracy trust only 3.67/5. The study is small, local, and lacks a randomized direct-workflow control. Its enduring contribution is the visible correction surface: precise-looking output remains editable evidence.
ChartAgent gives GPT-4o more than 40 chart-specific tools for OCR, segmentation, color, geometry, statistics, and calculation. It runs a ReAct loop for up to 15 iterations. On ChartBench it reached 71.39 overall versus 54.53 for GPT-4o alone. Adding the system to GPT-4o improved unannotated numeric questions by 26.7 points. Chart-specific image and analysis tools beat generic image tools by 30.0 points overall.
This is unusually useful ablation evidence: it compares tool roles around the same orchestrator. It also exposes the cost of composition. In 30 inspected trajectories, tool output was correct without recovery half the time. Of the rest, 70% recovered and 30% failed—15% of all inspected trajectories. Most failures were perceptual: obstructed text, poor contrast, occlusion, segmentation, overlap, and axis problems. The paper studies single-chart question answering, not authoring, dashboards, or reader outcomes.
Making multimodal LLMs reliable chart data extractors attacks a different failure: repeated calls to the same model disagree. Across 20 TinyChart samples, 99.4% of 1,000 charts produced differing outputs and 49% produced different table dimensions. The authors align repeated tables and take per-cell medians, stopping early when results converge and using disagreement as uncertainty. This improves several open and specialist models. It costs 4–16 times more inference, and uncertainty correlates only modestly with accuracy. The method helps triage unstable cells; it does not certify them.
Current reading: the useful intermediate is plural. Keep source data, chart specification, extracted table, rendered image, localized marks, and tool observations available for different checks. A pipeline that converts the chart to one table and discards the image makes some questions easier by making others impossible.
Total-workflow cost: the pieces do not yet add up
The literature can now supply cost units, but not a specialist-versus-general total. The studies use different tasks, denominators, acceptance rules, model versions, and price schedules. Adding one paper’s model bill to another paper’s human study would create a number without a workflow behind it.
| Evidence fragment | Reported unit | Missing from the same denominator |
|---|---|---|
| Self-ensembled extraction | About 4–16× API inference cost; up to 20 samples per image; every reported ensembled API configuration under $15 for a full benchmark | human verification, evaluator work, local compute, delivery, maintenance, readers |
| ExChart correction | 12 people, 24 charts each; mean 31.35 seconds for six-point charts and 32.87 seconds for fifteen-point charts, with pixel-level completion enforced | same-protocol current-general route, machine cost/latency, abandonment, maintenance, delivery |
| ChartAgent | 5–7 average iterations; about 6–10 seconds for one GPT-4o call versus about 90 seconds serial or 30 seconds parallel; about $0.40 per query across 4,952 pairs | human verification, evaluator labor, accepted-answer cost, maintenance, delivery, readers |
| ChartAgent failure audit | 15% unresolved tool-level failures in 30 sampled trajectories; fallback below 10% in a separate 30-trajectory review | larger denominator, human repair, consequential escapes, abandonment |
| VisJudge-Bench evaluation | Three crowd ratings per each of 3,090 samples and three experts reviewing all samples; 11% overall and 16% sub-dimension adjustment | model cost/latency, creator repair, expert time/pay, maintenance, readers |
| DashboardMimic evaluation | 450 manually annotated interaction tasks and three experienced evaluators scoring 90 dashboards | evaluator time/pay, generation and judge cost, correction, maintenance, mobile/accessibility/readers |
| Chart2Code-MoLA preparation and inference | Reports shared-corpus training memory/time and per-chart latency for chart-specific MoE + LoRA, full-fine-tuning, and LoRA-only variants | current-general route, consistent denominator and selected configuration, raw telemetry, human acceptance/correction, delivery, readers, maintenance; exact numbers quarantined after the audit below |
| METAL iterative chart reconstruction | Same-task direct, Best-of-N, and four-role iterative results with two base models; up to five recurrences, a 512–8,192 token-budget plot, and local GPU-memory description | matched calls/tokens/compute, actual early-stop and retry distribution, latency/charges/evaluator compute, human acceptance/correction, delivery, readers, maintenance |
| ProMCP phase profiling | Adjacent six-stage token and latency attribution across three MCP deployment topologies, with initialization and tool discovery separated | visualization task and specialist comparator, released task traces and historical run manifest, charges/compute, human acceptance/correction, delivery, readers, maintenance |
| ChartAgent tool-integrated reasoning | Same-task ChartBench scheduler and component sweeps report 73.2–78.8% accuracy beside 3.0–9.8 mean tool invocations | exact ablation-task count, model calls/tokens, local tool compute, reflection, latency, retries, charges, evaluator/human work, and accepted/rejected/abandoned outcomes |
| VisCoder2 conditional self-debug | On 888 VisPlotBench tasks, VisCoder2-32B and GPT-4.1 both finish with 732 execution passes; the released round totals reconstruct to 584 versus 714 revision generations | comparable tokens, accelerator/provider use, latency, charges, evaluator work, human or production acceptance, delivery, and maintenance |
VisJudge-Bench permits one transparent derived range: 3,090 samples × three annotations ÷ 15 images per batch gives 618 batch-annotations. At the reported 30–60 minutes per batch and estimated $10/hour, that is approximately 309–618 hours and $3,090–$6,180 on the stated crowd-pay basis. It excludes platform fees, screening, discarded work, and all expert review. It is an evaluation fragment, not the benchmark’s total cost. Likewise, a local model’s $0 API charge is not zero compute, and ChartAgent’s cheaper GPT-4o-mini figure is a projection rather than a measured replacement run.
A defensible comparison needs one frozen task sample and acceptance rule, the current general route and specialist route under equal total budgets, and eleven visible lanes: route preparation/ownership, task/acceptance, success, inference, latency, evaluation, human correction, unresolved failures, maintenance, delivery, and reader outcomes. Keep five cost classes separate: fixed preparation, periodic ownership, marginal attempts, failure-contingent work, and downstream delivery or reader work. Declare a volume and time horizon before amortizing setup. Missing values stay null. Measured, recomputed, projected, allocated, and source-reported costs stay distinct. Report cost per eligible task, attempted task, accepted artifact, and successful reader task separately so abandonment cannot make a route look artificially cheap.
Tool calls and accuracy are not cost per accepted artifact
The new ChartAgent preprint is closer than most studies to a same-task cost-quality frontier. Its full route reports 78.6% accuracy with 6.1 mean tool invocations; removing the information-gain scheduler reports 77.9% with 8.4, removing GroupTalk 76.8% with 5.6, and removing the tool library 75.4% without a call count. The paper also varies its cost weight and cumulative budget. This supports a bounded mechanism claim: scheduler choices can move quality and one resource proxy together.
It does not price the route. Tool invocations omit the model work that selects them, local vision-model compute, final multi-expert reflection, tokens, latency, retries, evaluator work, and human correction. The exact ablation denominator is unstated, and aggregate accuracy does not identify artifacts that passed a declared acceptance contract.
Ledger v7 now holds twelve cost-relevant fragments. Zero has both equivalent observed per-arm route-wide cost and a frozen accepted-artifact denominator. It therefore defines the ratio explicitly:
reconciled route cost for the declared population and window / accepted artifact count
Rejected, abandoned, quarantined, and no-output attempts stay in the numerator. Only artifacts passing the frozen human or production contract enter the denominator. The ratio is undefined when none pass. Report eligible, attempted, candidate, accepted-without-repair, accepted-after-repair, rejected, abandoned, quarantined, no-output, delivered, and reader-successful counts beside separate per-unit ratios. Tool calls, per-attempt API averages, benchmark accuracy or F1, finish signals, and one-arm telemetry are not substitutes.
The same execution endpoint can hide different retry burden
VisCoder2, accepted at ICLR 2026, supplies a more informative retry record than a single final pass rate. Its VisPlotBench evaluation contains 888 tasks across eight visualization languages. The VisCoder2-32B route begins with 649 execution passes and GPT-4.1 with 563. After up to three conditional self-debug rounds, both finish at 732 of 888, or 82.4%.
The endpoint ties. The paths do not:
| Route | Initial execution pass | Revision attempts by round | New passes by round | Total revisions | Total generations |
|---|---|---|---|---|---|
| VisCoder2-32B | 649 / 888 | 239 · 184 · 161 | 55 · 23 · 5 | 584 | 1,472 |
| GPT-4.1 | 563 / 888 | 325 · 217 · 172 | 108 · 45 · 16 | 714 | 1,602 |
The paper’s aggregate tables therefore imply 130 fewer revision generations for the VisCoder2 route. They also expose a declining conditional rescue rate: 23.0%, 12.5%, and 3.1% for VisCoder2, versus 33.2%, 20.7%, and 9.3% for GPT-4.1. That makes retry burden a survival process: each round spends work on the tasks that have survived all earlier failures. A final pass rate or fixed maximum round count erases both the number of attempts and their diminishing return.
This is not a cost winner. The two routes use different models, and the paper does not report comparable task-level tokens, accelerator or provider use, latency, charges, evaluator work, or human correction. The commit-pinned implementation retains raw initial responses, but its OpenAI self-debug path prints usage and returns only response text; its local-model debug path likewise returns text without generation metadata. Judge records keep scores rather than usage or latency. Training is described as three epochs on eight H100 GPUs, without elapsed hours, energy, allocation, or labor. The released repository also does not carry the historical result bundle used for the paper tables.
Execution is only the paper’s mechanical gate. The paper itself includes cases where execution improves without corresponding visual-quality improvement. Neither route’s 732 passing programs were classified under a frozen human or production acceptance contract, delivered to intended readers, or followed through maintenance. The defensible result is therefore: same execution endpoint, different retry topology, zero accepted-cost comparison, and no route winner. A valid reopening receipt would join every task and retry to observed resources and then to accepted, rejected, abandoned, quarantined, no-output, delivered, and reader-successful states under one contract.
A matched inference budget is a real result—and still not total cost
The peer-reviewed Selective Test-Time Scaling visual-insights study is the strongest held fixed-budget near miss. It sends raw tabular data through profiling, visualization, chart checking, insight generation, pruning, and final report scoring. Across its main VIS Publication experiment, the baseline and four pruning policies use the same models and environment while total declared LLM calls stay within 2.3% of the baseline. A separate run keeps reported output tokens within 1.1%. At pruning ratio 0.6, the mean proxy-judge score rises from 61.64 to 65.86 while the number of generated final reports falls from 1,435 to 578.
That supports a narrow mechanism result: under one pipeline and declared partial inference budget, early pruning can allocate search more effectively. It does not compare a specialist route with a direct current-general route, and “final report” means a generated candidate rather than work accepted under a human or production contract.
The official implementation also keeps the cost boundary visible. Stage counters are assumptions rather than reconciled provider records; plot execution can call a code rectifier without adding that call to the budget; metadata failures return zero budget; errored downstream workers are excluded from aggregate totals; and token records contain completion tokens only. The repository describes local result manifests but does not ship the experiment traces at the audited commit. These are release-accounting limits, not evidence that the private experiment failed.
Four expert annotators calibrate the proxy judge on score-stratified reports from two datasets. That is stronger than an unvalidated automated rating, but it still does not classify every task as accepted, rejected, abandoned, quarantined, or no-output. DV-World adds a separate same-harness interaction- charge versus capability view. Neither fragment supplies the full numerator or denominator.
The operational distinction is now three receipts, not one:
| Receipt | What it can establish | What it cannot establish alone |
|---|---|---|
| Matched declared partial budget | Two policies were allotted closely matched calls, output tokens, or another named slice. | Actual all-resource equality, failure cost, people, or delivery. |
| Observed route-wide use | Every eligible task retains actual machine, evaluator, human, failure, preparation, delivery, and maintenance events. | How many outputs passed the real acceptance bar. |
| Cost per accepted artifact | Reconciled route-wide cost is divided by artifacts passing one frozen contract. | Reader success or later maintainability unless those are inside the contract and window. |
Practitioners should keep repair and discarded candidates beside the task. BI leaders should ask vendors for all three receipts. Researchers should release task traces and preserve failed branches. Newsrooms should define editorial, source, accessibility, and reader gates before timing. Product teams should instrument request and retry identity rather than increment assumed stage counters. Educators and accessibility specialists should put representative use inside acceptance. Sponsors should read the shortest defensible answer as: one useful fixed-budget allocation result, zero of twelve total-cost-per- accepted-artifact comparisons, and no route winner.
Equal rounds are not an equal budget
METAL is useful same-task mechanism evidence. At five recurrences, its authors report average automatic F1 of 51.78% versus 40.45% for direct prompting and 43.13% for Best-of-N with Llama 3.2-11B; with GPT-4o, the corresponding values are 86.46%, 81.26%, and 82.32%. The system separates generation, visual critique, code critique, revision, deterministic verification, and automatic evaluation. The authors also say explicitly that it costs more than direct prompting.
The matched n = 5 rows are not an equal-resource comparison. In the official
implementation at commit
41832c1, Best-of-N performs five generation
calls. The METAL path performs one initial generation plus visual-critique,
text-critique, and revision calls in each recurrence—up to 16 model calls at
five recurrences, plus a possible retry. The paper describes four calls per
iteration, so a historical run receipt is needed to reconcile the paper and
code definitions. The public wrapper and logs do not retain per-arm provider
request IDs, input/output token use, call timestamps, latency, charges,
accelerator-seconds, energy, or evaluator compute. The public runners also
recommend one process for METAL and eight for baselines, preventing elapsed
time from being compared without observed concurrency and occupancy.
This does not invalidate METAL’s automatic performance result. It means the result cannot establish equal-budget efficacy or total cost. The corrected comparison has two resource records:
| Receipt | Required content | Decision use |
|---|---|---|
| Planned cap | Calls/retries, tokens and image units, tool/evaluator work, accelerator identity/time, latency and concurrency, charges, human minutes, and fixed-to-downstream allocations | Freeze the resource rule before quality is known. |
| Observed use | Route/task/attempt/event IDs, timestamps, stop reason, actual use in every declared dimension, value status, and source receipt | Distinguish equal resources, unequal-but-costed routes, missing use, and conflicting use. |
Matching iterations, candidates, provider rates, nominal token caps, or model labels alone satisfies neither receipt. A provider rate without observed usage is a price schedule, not workflow cost.
A machine total still hides the bottleneck
ProMCP is adjacent evidence about tool-mediated agents, not a chart-specific cost comparison. Across 155 MCP-Bench and MCP-Universe tasks, its authors attribute tokens and latency to user context, planning, tool calls, tool results, context updates, and final synthesis; they record session initialization and tool discovery separately. Customized-client configurations concentrate much of their measured burden in planning and schema injection, while the off-the-shelf Claude Desktop configuration concentrates most of its MCP-Bench latency in final-answer synthesis. The point is structural: changing the host, client, model placement, context policy, or streaming behavior can move the expensive phase even when the visible workflow still looks like “model plus tools.”
The limits are equally important. The off-the-shelf path is reconstructed from exported conversation timestamps and cannot expose hidden retries. The study uses one workstation and lightweight-to-moderate tools, so database, remote data, rendering, OCR, or local vision work can produce a different bottleneck. It reports tokens and latency rather than accepted visualization work or total human and machine cost.
The official release at commit 524b4ad provides useful event fields but no
task-level traces or historical effective-settings manifest from which the
paper tables can be recomputed. Its advertised benchmark entry point imports a
benchmark.runner module that is not present. Checked-in defaults also differ
from the paper settings on rounds, caching, and concurrency. These findings do
not disprove the paper results; they prevent the public repository from being
treated as a reproduced run receipt.
Ledger v7 therefore retains one topology-bound event record beneath the existing workflow lanes. It retains cold or warm state, schema-set hash, host/client/ server/model/tool identity, transport, cache, streaming, concurrency, timeout and retry policy; stage timestamps, tokens and payload footprint; provider or accelerator receipt; stop reason, outcome, value status, and source receipt. Hidden phases remain missing rather than being inferred from user-visible messages. Phase telemetry makes the machine portion inspectable. It does not replace human correction, accepted delivery, accessible use, reader outcome, or maintenance evidence.
A cost table still needs an integrity gate
Chart2Code-MoLA is useful because it separates training resources from
per-chart inference on one synthetic chart-to-code setup. It is not a current-
general comparison, and a direct audit found four conflicts in the paper’s own
record. Its five displayed split rows sum to 110,000 training, 17,000
validation, and 18,000 test examples, while the total row says
112,000/24,000/24,000. A 0.45-versus-0.40-second latency difference is 12.5%,
not the stated 5%. The peak-memory table prints an unexplained 4.84 delta for
12.3 versus 15.0 GB. The implementation section specifies top-2 routing,
attention + MLP, and alpha 16, while the ablation/conclusion names
probabilistic routing, attention + output, and alpha 32.
Those conflicts are present in the original PDF, not introduced by text extraction. They do not show that the method fails. They do mean its exact success and efficiency numbers cannot enter a route ranking until a corrected record or reproducible receipt binds one data denominator and configuration. The ledger now checks route identity, denominator arithmetic, percentages, configuration binding, and value status before admitting telemetry. Passing that preflight makes a receipt usable; it does not establish efficacy or a winner.
For practitioners, time the path from route preparation through accepted delivery and keep cold start, discovery, planning, tool/result transfer, context update, and synthesis separate. For BI and analytics leaders, require a topology card and phase receipt beside the total so the pilot names what is local, cloud, cached, streamed, concurrent, or hidden. For editors, source checking, correction, mobile and keyboard proof, and publication belong in the ledger. For researchers, release task traces, effective run manifests, observability state, denominator recomputations, and a later-change receipt. For product and engineering teams, instrument retry edges, effective settings, usage and charges before optimizing. For learners and accessibility specialists, keep assistive-path preparation and reader use separate from creator speed. For sponsors, the short answer is still that no cost winner is known.
The comprehensive audit and machine-readable ledger remain in private research custody and make that comparison answerable. They are a protocol, not a completed comparison: no route ran, no budget was authorized, and no cost winner exists.
Quality and integrity critics: direct progress, narrow authority
VisJudge is the closest thing in the current literature to a visualization-specific visual critic. Its benchmark contains 3,090 visualizations: 1,041 single charts, 1,024 multiple-chart compositions, and 1,025 dashboards. Three crowd annotators rated each sample; three visualization experts reviewed all annotations. The rubric adapts six dimensions—fidelity, semantic readability, insight discovery, design style, visual composition, and color harmony—to the chart.
The resulting 7B model, trained from Qwen 2.5-VL with supervised and reinforcement learning, reached .421 mean absolute error and .687 correlation on the held-out test. GPT-5 reached .553/.428, GPT-4o .610/.482, and Claude 4 Sonnet .622/.465. This is a meaningful specialist win on a directly relevant task.
It is not a final acceptance gate:
- The critic sees static images, not the source data, chart specification, interaction, or responsive states. “Fidelity” therefore means what can be inferred from the presentation, not agreement with the source.
- Its correlation falls from .577 on single charts and .565 on multiple charts to .375 on dashboards.
- The crowd pool was 88.7% US-based, and a point rating compresses legitimate preference differences.
- When its feedback improved charts produced by seven generation models, the revised charts were evaluated automatically by GPT-4o. No human or reader outcome was measured.
- Generic Qwen feedback often made outputs worse. “Add a critic” is not the tested intervention; critic knowledge and rubric fit matter.
Misviz measures a narrower but important kind of critique: detecting 12 misleading-design categories. It includes 2,604 real-world visualizations and 57,665 synthetic charts. On real charts, GPT-o3 led at 80.0 F1 and 58.8 exact match; the image-plus-extracted-axis classifier reached 60.9 F1 and 12.3 exact match. On controlled synthetic charts, the rule-based and classifier systems became competitive. A linter with ground truth axes achieved 99.7 precision but only 52.2 recall. The real-world result favors broad general reasoning; the synthetic result favors encoded rules. Neither captures narrative deception, domain knowledge, or the full published taxonomy of misleading techniques.
VisDeception uses 1,600 paired charts: the same data rendered faithfully and with one of eight deceptive tactics. A multi-stage system first extracts structured metadata and then reasons about the chart. It reduced Gemini 2.5 Pro’s paired error from .099 to .024 and GPT-4o’s from .502 to .255. It worsened Gemini 2.5 Flash from .109 to .135 and Qwen3-VL-2B from .251 to .275. The generated paired design isolates causality well; it does not represent messy surrounding prose, captions, or real-world author intent.
Current reading: use a specialist critic as one bounded witness. Pair a perceptual-quality rubric with deterministic data/spec tests, an integrity taxonomy, accessibility checks, and contextual review. Its output should be a diagnosis with evidence, not a single universal score.
Grounding: the critic must point to what it means
A critique becomes more useful when it can locate the mark, label, axis, or legend that caused the judgment. ChartREG++ tests this directly with 850 chart images, 3,400 referring expressions, 18 element types, and point, box, and mask outputs. The references include textual, data, and visual cues.
Strong general multimodal models remained below 50 F1 for point and box grounding. For masks, a zero-shot SAM 3 candidate pipeline plus language alignment reached 25.01 and 21.84 F1 on the two evaluation sources. Synthetic chart-specific Mask2Former candidates plus the same alignment idea reached 52.55 and 49.85. This is evidence that generic open-vocabulary segmentation does not automatically resolve fine, overlapping chart marks. It is not a fair zero-shot comparison: the chart candidate model was trained specifically for the domain.
The chart-specific masks transferred to a modified and verified real-chart set, especially for lines, but the benchmark remains static and concentrated on common Matplotlib-style elements. Candidate quality and language alignment are still separate bottlenecks.
Current reading: grounding is a strong candidate for the critique stack because it can connect a finding to visible evidence and support a repair check. The missing experiment is not another mask score. It is whether grounding helps a system identify a real defect, edit the right object, and verify that the edit fixed the defect without changing something else.
OCR and document models: useful sensors, not quality judges
Three recent releases matter because many real visualizations arrive inside reports, dashboards, slide decks, and screenshots rather than as clean chart images.
- PaddleOCR-VL-1.5 is a 0.9B model for layout segmentation and element recognition across text, tables, formulas, charts, and reading order. It reports 94.50 overall on OmniDocBench v1.5 versus 90.33 for Gemini 3 Pro and 85.50 for GPT-5.2 under the paper’s evaluation. Some current comparator results were run independently; the team also created the Real5 degradation benchmark.
- HunyuanOCR-1.5 is a 1B model with 4K input, 128K output context, chart parsing, and an accelerated decoder. It reports 48.9 English and 64.1 Chinese overall on ChartArena, above dots.mocr and PaddleOCR-VL-1.6 in the same table, and reduces one vLLM page latency from 3.032 to 1.408 seconds with its faster decoder. Its best hallucination-page recall on CHAOS is only 14.15, a reminder that relative leadership can coexist with a low absolute result.
- dots.mocr is a 3B OCR model with optional SVG-style structured graphics output. It reports lower text-edit and reading-order error than Gemini 2.5 and Gemini 3 on OmniDocBench v1.5, and 83.2 on ChartQA. Some structured-graphics evaluation uses an automatic model judge, and the model card acknowledges that tables, formulas, and graphics are not fully robust.
These systems may recover labels, reading order, chart regions, or a structured graphic before a critic reasons. That is a plausible architecture role, not a demonstrated outcome. None of these studies compares the same visualization critic with and without the parser, measures whether source-data errors are found, or tests comprehension after the parsed evidence is used.
Current reading: add a document specialist when OCR or layout is a measured bottleneck. Evaluate the entire downstream task. Do not convert a document benchmark lead into authority over chart quality.
General-model capability is improving, but not uniformly
Two human-comparison studies help locate the moving baseline.
CHART-6 adapted six independently developed human visualization-literacy tests to eight vision-language models. The strongest 2024-era model, GPT-4V, remained below human performance on several tests, and model error patterns were far from the human noise ceiling. The work also found that invalid answer formatting materially changed scores—a harness issue that can masquerade as visual reasoning.
A July 2026 scientific visualization literacy study tested six newer models on 49 items spanning 18 scientific visualizations and illustrations, with 485 human nonexperts. Gemini 3.1 Pro Preview reached 88.6% overall; Claude Opus 4.6 reached 75.3, GPT-5.4 75.1, and the human table reports 75.6. The paper’s narrative gives a slightly different 75.9 human figure. The result does not mean Gemini has universal scientific visualization literacy: some task categories contain only one or two items, and fine quantitative reading, flow direction, and unsupported encodings remain difficult.
The baseline has plainly moved. What has not disappeared is the need for specialized evidence, executable checks, and task-specific evaluation. Better general models make generic “chart expertise” packages less defensible; they make narrow specialists with inspectable added information more valuable.
Chartography makes the remaining professional reading gap concrete. Its 100 deliberately difficult tasks were authored by practitioners across 12 domain labels and independently verified by three experts each. The best of 30 frontier configurations reached 45.0% mean pass@1. More reasoning helped most paired configurations by a median 4.5 percentage points, but often extended an initial misreading of sparse axes, 3D projection, contours, or domain conventions. Because the benchmark screened for difficulty, the result is not a failure rate over ordinary charts. It is evidence that professional environment and visual form remain routing variables.
Stateful dashboards: a screenshot is not the whole artifact
Dashboard2Code adds 180 Plotly Dash dashboard-code pairs, 20 visualization types, eight callback patterns, and 450 interaction tasks. The best reported configuration scored 79.4 overall and 64.2 on the hardest interaction level. DOM access raised one model’s successful UI exploration from 20.61% to 64.58%. Ninety generated dashboards were also scored by three visualization-experienced graduate evaluators; the final automatic metric correlated .781 with those human ratings.
The benchmark exposes failures a static critic misses: hidden state, cross-control dependencies, and factually wrong transformations can coexist with a visually responsive interface. It also draws a firm scope boundary: Plotly Dash only, fixed 1920×1080 viewport, no animation or popups, and no mobile, keyboard, assistive-technology, or reader evaluation.
Context changes what “better” means
A journalism graphic, a BI dashboard, a scientific figure, and an accessible description do not share one quality function.
- Journalism and explanation require claim fidelity, annotation, source context, mobile reading, and reader comprehension. A quality critic can flag composition; it cannot validate the reporting or the reader’s takeaway.
- BI and operational dashboards require governed definitions, current state, permissions, interaction, responsive behavior, and decision support. Static chart scores omit much of the product.
- Scientific visualization requires domain conventions, unfamiliar encodings, uncertainty, and sometimes animation or 3D spatial reasoning. The 2026 literacy study shows strong progress and uneven task performance.
- Document recovery prioritizes text, layout, table, formula, and reading order. OCR specialists are directly relevant here, even if they have no aesthetic judgment.
- Accessibility requires structure and descriptions that help actual blind and low-vision users complete tasks. A 2026 description study generated 1,224 outputs and 17,755 claims; 73.9% of claims in “Real-World Significance” were labeled speculative. Adding accessible context improved direct numeric agreement for some models, while adding the image to full context produced no consistent numeric benefit. The provenance labels were themselves model-assigned, and no blind or low-vision users were studied. A separate 12-participant BLV study does supply direct learner experience: eleven preferred tactile charts plus text and an LLM, one said the better mode depended on complexity, and none preferred text and chat alone. Participants described touch as supplying spatial structure and chat as supplying flexible clarification, but measured chart-understanding accuracy did not improve. Preference, mental model, and comprehension are separate acceptance outcomes.
The common spine is the same—intent, source, representation, rendering, verification, delivery, reader—but the specialist and acceptance evidence must be delegated by context.
Where the literature is converging
Preserve structure and pixels
Tables, chart specifications, OCR blocks, masks, and images are complementary. The recurring architecture exposes intermediate evidence instead of hiding all perception and reasoning inside one completion.
Separate perception from judgment
The papers increasingly isolate axis recovery, mark localization, calculation, deception detection, and perceptual rating. This makes failures diagnosable and allows a system to assign the smallest relevant specialist.
Use executable operations for calculations and checks
Program-of-thought, chart analysis tools, linters, and deterministic source/spec validation reduce reliance on visual arithmetic. They remain bounded by their inputs and taxonomies.
Keep correction mixed-initiative
ExChart’s overlay, ChartAgent’s visible tools, and ensemble disagreement all make uncertainty inspectable. They support a human deciding whether and how to repair rather than merely accepting a confident answer.
Test on harder distributions
ChartQAPro, CharXiv, Misviz’s real split, ChartREG++, scientific visualization literacy tests, and accessibility-context ablations all move beyond clean synthetic single charts. The hardest findings are often transfer failures, not another point of in-distribution progress.
Where approaches legitimately diverge
- End-to-end general model versus decomposed pipeline. General models lead on heterogeneous real charts and broad context. Decomposed systems lead when a narrow operation—axis calibration, mask generation, calculation, or a quality rubric—adds information the model lacks.
- Learned critic versus deterministic rule. Learned critics cover fuzzy composition and insight dimensions. Rules provide high precision for a bounded integrity condition. The right choice depends on whether false positives or missed defects are more costly.
- Table versus visual grounding. A table is strong for values and weak for visual encodings. Grounding preserves spatial evidence but does not recover a trustworthy data table by itself. Many systems need both.
- Single pass versus ensemble. One pass minimizes latency and cost. Repetition helps when extraction instability is common and errors are costly, provided the system has a stopping and escalation rule.
- Compact local specialist versus frontier API. Small models can win a narrow same-test task and offer privacy or latency benefits. Frontier models can transfer better across unfamiliar visuals. Deployment constraints and the actual held-out task should decide.
What has performed poorly or caused harm
These are not theoretical warnings; each has direct evidence in the reviewed work.
| Technique or assumption | Negative finding | Practical consequence |
|---|---|---|
| Select a model because it led an older chart benchmark | ChartGemma and TinyChart collapse on ChartQAPro; DePlot and TinyChart trail frontier general models on WB-ChartExtract | Re-run same-task baselines after model and benchmark changes |
| Treat a table as a lossless chart representation | DePlot loses color, orientation, and visual attributes and trails MatCha on PlotQA | Preserve the rendered visual and explicit encodings alongside extracted values |
| Train on synthetic charts and assume real-world transfer | Misviz method rankings change sharply between synthetic and real charts | Maintain separate synthetic diagnostic and real acceptance sets |
| Add generic visual feedback | Generic Qwen feedback often lowers downstream chart quality in the VisJudge experiment | Evaluate the critic itself and expose its rubric |
| Add a vision caption by default | In three small product-owned Vizier comparisons, the captioned arm won 0/3, 1/5, and 1/5 cases while no-caption won 1, 3, and 4 | Treat captions as routed, fallible sensor output; retain them only when they improve a named downstream decision |
| Apply one mitigation pipeline to every model | VisDeception’s structured mitigation improves several models and worsens Gemini 2.5 Flash and Qwen3-VL-2B | Route and ablate by model/version; preserve a no-mitigation baseline |
| Use a general segmenter for fine chart marks | SAM 3 candidate grounding remains weak on ChartREG++ | Use chart-aware candidates when grounding is a measured bottleneck |
| Trust a single extraction | 99.4% of TinyChart cases varied across 20 samples; 49% changed table dimensions | Surface instability and require correction for consequential use |
| Treat ensemble disagreement as confidence | Correlation with accuracy is only about −.30 to −.37 and inference costs grow 4–16× | Use it as escalation evidence, not a probability of correctness |
| Treat an automatic judge or preference as reader comprehension | VisJudge revisions are automatically evaluated; the BLV tactile + text + LLM condition was strongly preferred but did not improve measured accuracy | Add human expert and reader acceptance, and keep preference, mental model, comprehension, and decisions separate |
| Treat a fixed desktop dashboard as responsive evidence | Dashboard2Code measures stateful Plotly Dash behavior at 1920×1080 and excludes animation and popups | Test multiple viewports, responsive reflow, animation, popups, keyboard paths, and assistive technology |
A reference architecture for critique
The research supports a layered system with explicit evidence ownership:
| Layer | Primary evidence | Appropriate mechanism | Authority boundary |
|---|---|---|---|
| Intent and context | audience, task, claim, stakes, environment, device | human brief + general reasoner | determines which quality dimensions matter |
| Source truth | data, units, joins, transformations, semantic definitions | deterministic queries, tests, provenance | authoritative for values; no image model can replace it |
| Chart structure | specification, scales, encodings, labels, states | schema checks, linter, chart parser | validates declared structure, not reader interpretation |
| Rendered evidence | pixels, text, geometry, marks, occlusion | OCR, chart extraction, grounding/segmentation | identifies visible evidence and discrepancies |
| Task-specific critique | integrity taxonomy, perceptual rubric, domain conventions | specialist critic + deterministic rules | limited to tested dimensions and distributions |
| Broad synthesis | surrounding prose, purpose, exceptions, tradeoffs | current general multimodal model | integrates context; should cite upstream evidence |
| Delivery | responsive states, interaction, keyboard, screen reader, latency | browser/device replay and accessibility tests | validates the artifact people receive |
| Acceptance | comprehension, decision quality, harm, editorial or domain judgment | intended readers and accountable humans | final authority for consequential communication |
This architecture does not require every layer for every chart. A simple local plot may need source checks and a rendered inspection. A public dashboard may need all eight. The important design choice is that no component silently gains authority over a layer it cannot observe.
Adoption signals: evidence of attention, not effectiveness
Several projects have public repositories. As captured on 2026-08-14:
| Project | GitHub stars | Forks | What it indicates |
|---|---|---|---|
| SAM 3 | 11,315 | 1,704 | substantial attention to a broad segmentation model, not chart-grounding quality |
| dots.ocr | 9,069 | 802 | strong interest in an open OCR/structured-document model |
| HunyuanOCR | 1,920 | 149 | material attention to a lightweight OCR family |
| VisJudgeBench | 124 | 6 | early research uptake for a new visualization critic |
| ChartQAPro | 45 | 7 | a benchmark repository, where stars are especially weak as use evidence |
Stars and forks can prioritize integration research. They do not establish downloads, active deployments, retention, task fit, reproducibility, or output quality. SAM 3’s broad adoption signal is particularly easy to misread: ChartREG++ directly finds that a generic SAM 3 candidate pipeline is not enough for fine chart-element grounding.
Gap ledger
| Question | Status | Evidence in hand | Evidence still needed |
|---|---|---|---|
| Do chart specialists beat strong general models on realistic OOD tasks? | Answered: no class-wide ordering | ChartQAPro and WB-ChartExtract show severe specialist transfer gaps; narrow author-benchmark wins also exist | Reopen after independently replicated dominance across realistic tasks |
| Do chart-specific tools add value? | Answered for chart QA | ChartAgent’s same-orchestrator ablations show large gains over no tools and generic tools | Replication on newer models, dashboards, generation, and critique |
| Can a quality critic improve final charts for people? | Partly answered | VisJudge improves expert-rating prediction and automatic downstream scores | Human expert and reader evaluation of the revisions |
| Does stronger OCR/document parsing improve critique? | Open | PaddleOCR-VL, HunyuanOCR, and dots.mocr improve upstream recovery | End-to-end ablation with the same critic and acceptance task |
| Does chart grounding improve defect localization and repair verification? | Partly answered | ChartREG++ shows chart-aware masks improve localization | Diverse real charts and a repair-verification outcome |
| Can model disagreement drive safe escalation? | Partly answered | Repeated sampling improves extraction and reveals instability | Calibrated thresholds across models, distributions, costs, and stakes |
| Do specialist-assisted workflows improve reader or accessibility outcomes? | Partly answered | A 12-participant BLV study finds strong preference and spatial-model benefit for tactile + text + LLM, but no accuracy lift and no isolated chart-vision specialist | Delivered-chart specialist ablation on comprehension, decisions, mobile use, and assistive technology |
| What happens on interaction, animation, responsive states, and dashboards? | Partly answered | Dashboard2Code measures callbacks and hidden state in 180 fixed-desktop Plotly Dash applications; VisJudge dashboard correlation is weaker | Multi-viewport stateful benchmark with browser traces, responsive rendering, animation, keyboard paths, and assistive technology |
| Which specialist belongs in which environment? | Partly answered | The role typology explains why routing should differ; Chartography shows professional domain conventions and visual forms materially affect performance | Common journalism, BI, science, operations, and accessibility routing comparison |
| What is the total human and machine cost? | Partly answered: ledger corrected, comparison not run | Twelve fragments expose partial preparation, inference, latency, correction, evaluator, failure, call-topology, phase, interaction-charge, retry-survival, and tool-use units; one matches a declared partial inference budget, while zero has both equivalent observed per-arm route cost and a frozen accepted-artifact denominator; ledger v7 keeps the three receipt levels separate | Run current-general and specialist routes on the same tasks with one acceptance rule, equivalent observed usage, all outcome states, preparation and amortization, delivery, reader task, and later maintenance event |
| Can the held studies support cost per accepted artifact? | Answered: zero of twelve held comparisons | Selective TTS is the strongest fixed-budget near miss; VisCoder2 adds a reconstructable same-endpoint retry topology, but neither supplies equivalent observed route use plus task-level acceptance | Reopen after a held study or authorized run reports equivalent arm receipts and accepted, rejected, abandoned, quarantined, and no-output states under one contract |
This ledger is the practical boundary of the review. The first two questions have bounded answers. Six have meaningful component or adjacent evidence but remain open at the workflow or reader level. Two remain direct corpus gaps.
Recommended experiments, in order
1. Establish the moving baseline
Use the current general multimodal model with its normal harness. Measure the actual local task, include hard and unanswerable cases, and record value accuracy, defect detection, latency, cost, and human correction. Re-run after a material model or harness release.
2. Add deterministic source and specification checks
Before another model, verify values, units, joins, transformations, scales, labels, and expected states from authoritative inputs. This establishes which errors actually require vision.
3. Add one perception specialist at a measured bottleneck
If OCR fails, test a document parser. If values cannot be recovered, test a chart extractor. If a critique cannot point to the right mark, test chart grounding. Keep the rest of the system fixed and score the end-to-end task, not the component’s published benchmark.
4. Compare critic roles separately
Run a perceptual-quality critic, an integrity taxonomy, deterministic rules, and the general model on the same charts. Measure agreement with accountable experts by dimension. Do not collapse the outputs into one score until each component’s false-positive and false-negative costs are known.
5. Test repair, not only detection
Require the system to localize a defect, propose or perform the smallest edit, and prove that the targeted defect improved without regression. Compare a grounded critic with a text-only critic and preserve the original artifact.
6. Evaluate ensemble routing under a fixed budget
Compare single pass, repeated sampling, and escalation to a stronger model or human. Set one total inference, latency, and reviewer-time budget. Evaluate calibration and abandoned work, not only average extraction gain.
7. Finish at the delivered reader experience
Test desktop and phone layouts, interaction states, keyboard access, screen reader output where relevant, comprehension, confidence, and decision quality. A better screenshot score is an intermediate result, not the goal of data visualization.
Method and limits
This is a purposive deep reading of 26 papers plus primary model cards, repository records, and three product-owned historical caption-comparison receipts. The local receipts are self-reported, small, and kept separate from independent research. The review emphasizes recent 2025–2026 work while retaining the 2023–2024 systems needed to explain the technique history. Papers were read in full where readable text was captured. Numerical comparisons appear only when reported on the same test and metric. Author-benchmark results are labeled; automatic judges, small human samples, preprint status, and static/synthetic scope are carried into the interpretation.
The review is not exhaustive. Most evidence is English-language, static, and same-team. Frontier models move faster than peer review and are not evaluated on every specialist benchmark. Repository popularity is not usage. No local deployment, cost, privacy, licensing, or hardware evaluation was performed.
The central conclusion is therefore intentionally conditional: route a specialist when it supplies a tested piece of evidence that the baseline lacks, and keep its authority no broader than that evidence.
Update log
-
2026-08-16 — Conditional retry burden separated from cost. Recomputed VisCoder2 and GPT-4.1 self-debug paths over 888 tasks: both finish at 732 execution passes, while the routes use 584 versus 714 revisions. The released debug paths do not retain comparable resource telemetry, and execution is not human or production acceptance. Ledger v7 therefore advances to twelve held fragments, 0/12 accepted-cost comparisons, and no route winner.
-
2026-08-15 — Matched inference budget separated from total accepted cost. Added the Selective TTS visual-insights paper and commit-pinned code audit, plus DV-World’s same-harness interaction-charge view. One of eleven held fragments now matches a declared partial call/output-token budget; zero has equivalent observed route-wide use plus frozen accepted outcomes. Ledger v6 keeps those receipt levels separate and selects no winner.
- 2026-08-15 — Accepted-artifact denominator made explicit. Added ChartAgent’s same-task accuracy/tool-call sweeps as a partial cost proxy, found zero of nine held fragments with both equivalent route cost and a frozen accepted-artifact denominator, and advanced the ledger to v5. No route ran and no cost winner exists.
- 2026-08-15 — Machine totals split by phase and topology. Added adjacent ProMCP evidence and a commit-pinned release audit. Ledger v4 now preserves cold start, discovery, execution phase, effective topology and runtime state; the source does not add a visualization comparison, reproduced run, or cost winner.
- 2026-08-15 — Equal rounds separated from equal resources. Added METAL’s
same-task chart-reconstruction result, audited its paper and official code,
and found that matched
n = 5rows do not match calls or retain observed usage. Ledger v3 now requires planned and observed receipts over twelve resource dimensions. No route ran and no cost winner exists. - 2026-08-15 — Cost starts before inference; source numbers must reconcile. Added one peer-reviewed preparation/inference fragment, quarantined its exact figures after denominator, percentage, table-delta, and configuration conflicts, and expanded the protocol to eleven lanes, five cost classes, and a source-integrity preflight. No route ran and no winner exists.
- 2026-08-15 — Total-cost comparison made measurable, not measured. Audited five primary evidence fragments, exposed inference, latency, correction, evaluation, and unresolved-failure units, and added a ten-lane equal-budget ledger with explicit missingness. No route ran and no total or winner exists.
- 2026-08-14 — Initial public edition. Compared six specialist vision roles, recent models and benchmarks, same-test gains, transfer failures, critic authority, and routing experiments.