Source map

Every score is a synthesis. These are the studies closest to each judgment.

The links beside current scores and forecast tests take readers directly to the relevant evidence. This register explains what each evidence block supports here—and what claim would overreach it.

Evidence blockPrimary sourcesSupports hereDoes not support

Definition of the ideal

Separate purpose, abstraction, representation, interaction, algorithm, work practice, user performance, and human outcome before selecting an evaluation.

The nine dimensions or their maturity scores; those are this report’s synthesis.

2017–2020 anchors

Declarative interaction, learned generation, constraint-based recommendation, and natural-language specifications had become demonstrable and inspectable.

Reliable framing, real-world delivery, or reader benefit.

2022–2023 anchors

Real-chart question answering and multi-stage LLM visualization generation crossed into repeatable public evaluation.

Professional reading reliability, rendered-chart human quality, or downstream outcomes.

Typed interactive-document authoring

A human-readable State–Render–Transition–Constraint plan can expose interaction intent and improve same-pipeline authoring measures.

Independent replication, source truth, representative readers, accessibility, production delivery, maintenance, or autonomy.

Human outcome and lifecycle

Bounded comprehension, preference, accessibility, novice production, correction, accepted second-person change, one repeat contributor, one project-authored recovery, one open proposal, and one multi-attempt chain through maintainer recheck can be studied directly.

Representative professional reader benefit, maintenance-authority transfer, or a complete production-to-maintenance episode.