The same scorecard through time
Progress accelerated after 2023, but no common-core dimension has crossed into field-wide deliverability.
The rubric is fixed at the August 2026 ideal. Dates mark changes in public evidence, not annual averages. Exact cells are editorial synthesis, not normalized benchmark scores.
Each lane repeats its dates so mobile readers do not need to remember a distant header.
Data, semantics, and provenance
Visual construction fidelity
Visual interpretation and reasoning
Integrity critique and uncertainty
Iterative steering and repair
Interaction, responsiveness, and accessibility
Reader and decision outcomes
Production efficiency, governance, and maintenance
The substrate became programmable.
Declarative grammars, constraints, learned generation, and natural-language specifications made parts of construction inspectable. Evidence remained example-led.
Understanding and LLM generation became repeatable.
Real-chart question answering and multi-stage LLM pipelines crossed into benchmarked capability, while outcome validation stayed absent.
Systems became steerable and executable.
Mixed initiative, render feedback, database execution, local edits, and coding harnesses moved several bounded workflows to usable maturity.
The frontier broadened sideways.
Professional, multilingual, multi-chart, document, state, replay, and judge-reliability tests expanded scope without closing the outcome and lifecycle gaps.