What the benchmarks actually measure

A benchmark result is a task-and-grader result, not a field-wide score.

The same system can pass code checks, misread the finished chart, miss a deceptive axis, or fail after a dashboard control changes state. Read the evidence as cumulative obligations. No captured evaluation spans the whole episode.

  1. 01

    Structure

    Requested code, fields, marks, layout, or specification

  2. 02

    Execution

    Compile, render, load, and callback completion

  3. 03

    Semantics

    Values, transformations, bindings, and encodings

  4. 04

    Perception

    Extraction, grounding, questions, and comparison

  5. 05

    Integrity

    Deception, defects, and calibrated quality judgment

  6. 06

    Interaction

    Control behavior, hidden state, and intended-use replay

  7. 07

    Human outcome

    Accepted work, comprehension, decisions, and access

Evaluation family What the system must do Representative evidence What it still cannot establish
Creation

Generation and reconstruction

Turn an image, source data, database question, or specified view into chart code, a visualization query, or a compiled artifact.

Plot2Code · VisEval · Text2Vis · RealChart2Code · Raiven

Execution and resemblance do not establish that the question, source, semantics, or design is appropriate for a person.

Reading

Perception, reasoning, and integrity

Extract values, locate marks, answer questions, detect deceptive encodings, or predict an expert-defined quality rating.

ChartQAPro · CHART-6 · Misviz · VisJudge · FinChart-Bench · Chartography

Chart reading is not chart authoring. A visual critic usually cannot verify raw data, audience fit, interaction, or reader outcome.

Context

Multiple charts, documents, and languages

Compare chart pairs, localize and reason across multi-panel figures, retrieve charts and text from documents, or work outside English.

ChartDiff · multi-chart PolyChartQA · Chart-MRAG · multilingual POLYCHARTQA

Translated, generated, or carefully curated sources do not represent ordinary work across languages, institutions, devices, or local chart conventions.

Use

Interaction and people

Navigate, reconstruct, or generate an interactive dashboard; replay intended use; or observe a creator or reader performing a task.

DashboardQA · Dashboard2Code · DashArena · bounded authoring and reader studies

No evaluation follows one artifact through authoritative data, acceptance, delivery, mobile and assistive use, reader outcome, later correction, and maintenance.

Question source changed difficulty

Up to 27.4 points lower

Multi-chart accuracy fell on human-authored questions compared with model-generated questions.

The metric changed the winner

ROUGE and human-aligned quality disagreed

ChartDiff specialists and pipelines scored higher on overlap while general models led on calibrated quality.

The intervention created regressions

+15.4–19.6 on target cases; as low as −8.5 elsewhere

Table-based QA resisted misleading charts but sometimes damaged accuracy on non-misleading charts.