What changed this month
DashArena tests whether a dashboard supports analysis—not whether a reader learns.
Its judge sees the task, schema, screenshots, and a browser-replayed interaction trajectory. Adding interaction evidence raised agreement with dashboard-experienced humans by 8.1 percentage points.
No benchmark measures communicative value.
Task-grounded analytical support has a credible measure. Reader comprehension, retention, and decision quality remain open.