The evidence ladder

“Works” is seven different claims.

Passing a lower layer is necessary. It is never a proxy for all the layers above it.

  1. 01

    Structure

    Required fields, files, or spec entries exist.

    AntV · OpenAI skill · Vizro not execution
  2. 02

    Execution

    Code runs and a nonblank artifact appears.

    Raiven · Vizro · DashArena not correct values
  3. 03

    Data fidelity

    Values, joins, filters, and aggregations match the source.

    nvAgent · Raiven · Vizro not legibility
  4. 04

    Visual integrity

    Encodings are coherent and not misleading.

    Raiven · PlotGen not insight
  5. 05

    Interaction

    Controls and linked views behave as intended.

    DashArena · Vizro not a good analysis
  6. 06

    Analytical support

    The artifact helps address the task.

    DashArena not learning or decisions
  7. 07

    Reader outcome

    A defined audience understands, remembers, or decides better.

    Open in this scan requires real readers
Evidence layers used throughout this review. Data Formulator 2 and DashChat add small human studies of authoring and prototyping experience; they do not measure the eventual reader’s comprehension or decision.

What changed this month

DashArena tests whether a dashboard supports analysis—not whether a reader learns.

Its judge sees the task, schema, screenshots, and a browser-replayed interaction trajectory. Adding interaction evidence raised agreement with dashboard-experienced humans by 8.1 percentage points.

Earlier wording

No benchmark measures communicative value.

Better wording now

Task-grounded analytical support has a credible measure. Reader comprehension, retention, and decision quality remain open.

Judge agreement with humans
Full evidence
79.8%
No interaction evidence
71.7%
Rules only
42.4%
DashArena · 100 labeled pairs; 99 evaluable · six dashboard-experienced annotators

Why a clean render is weak evidence

Failure survives at every layer.

No tested model exceeded 86% render success or 74% trajectory replay success. Among 30 execution-clean failures, 21 still contained semantic defects.