Benchmark · 2026

DashArena

A benchmark, browser executor, and human-calibrated judge for open-ended interactive dashboard generation.

Open the primary source ↗

Interactive dashboard generation

Generate a single-file dashboard and a two-turn intended-use trajectory for 234 open-ended tasks.

Browser render and replay, schema and execution receipts, and human-calibrated pairwise judging.

Adds task-grounded interaction evidence and shows that current systems can attempt analytical dashboards while remaining unreliable in render, replay, and semantics.

Model-authored trajectories and Tableau-derived tasks do not represent every audience, device, decision, or production environment.

Where this appears in the research.

Audience routes that point here.