Current common-core scorecard

The field is closest on bounded creation and farthest on what happens to people and artifacts afterward.

Each bar shows current maturity against the ideal target. “3” means usable somewhere under declared conditions—not reliable everywhere.

Current maturityRemaining distanceH / M Evidence confidence
DimensionCurrent → idealConfidenceBinding reason · evidence

Task and audience framing

123452 → 4
M

Analytic tasks can be inferred; the right purpose, audience, consequence, and authority usually remain supplied by people.Evidence: NL4DV-LLM ↗Data Formulator 2 ↗

Data, semantics, and provenance

123453 → 4
M

Usable grounding exists in governed and mixed-initiative settings; source choice, correction state, and full lineage remain inconsistent.Evidence: nvAgent / VisEval ↗Data Formulator 2 ↗

Visual construction fidelity

123453 → 4
H

Conventional static work is usable. Real-data, multi-turn, scientific, and interactive tasks still show material failures.Evidence: RealChart2Code ↗Raiven ↗DashArena ↗

Visual interpretation and reasoning

123453 → 4
H

Basic chart QA is strong; hard professional, multilingual, multi-chart, and document cases still expose perception and convention failures.Evidence: Chartography ↗POLYCHARTQA ↗Chart-MRAG ↗

Integrity critique and uncertainty

123452 → 4
H

Misleading-chart accuracy sits near random on one broad study; corrective representations help conditionally and can introduce new errors.Evidence: Misleading-chart interventions ↗Misviz ↗

Interaction, responsiveness, and accessibility

123452 → 4
H

Desktop state and replay are testable. Mobile, responsive, keyboard, assistive, authenticated, and production behavior are not evaluated together.Evidence: DashboardQA ↗Dashboard2Code ↗DashArena ↗

Production efficiency, governance, and maintenance

123451 → 4
M

Three dashboard afterlives expose different stopping points. Prism and OpenClaw also preserve accepted second-person AI-assisted change, Prism has a repeat contributor, and OpenClaw has an open repair proposal. None establishes maintenance authority or independent final acceptance. Whole cost, governance, and reader outcomes are not closed together.Evidence: KubeStellar ↗OpenClaw ↗Prism ↗