# How close is AI-assisted data visualization to the ideal?

Status: field-level capability scorecard, evidence cut 14 August 2026. Historical
scores are retrospective synthesis, not contemporaneous benchmark totals.
Forecast scores are conditional translations of previously published resolution
tests, not promises.

Companions: [The state of AI-assisted data visualization
research](/reports/research-review/), [What the
benchmarks actually
measure](/reports/benchmark-crosswalk/), [The
practitioner and reader
experience](/reports/practitioner-and-reader-experience/),
and [What would unlock the next
capabilities?](/reports/capability-forecast/).

## Executive answer

We can define an ideal state, but not as one universally best chart or one
autonomous agent.

The common ideal is a system that can:

1. understand the actual human purpose and audience;
2. ground its work in authoritative data, definitions, and provenance;
3. construct a faithful and appropriate visual artifact;
4. interpret visualizations accurately, including difficult and unfamiliar ones;
5. detect misleading or uncertain evidence and calibrate its confidence;
6. support local, non-regressive revision;
7. work in the delivered interactive, responsive, and accessible surface;
8. improve the intended human outcome; and
9. remain governable, maintainable, and worthwhile over its lifecycle.

The context determines what the last mile means. An exploratory system should
help an analyst form and test valid questions. A newsroom graphic should improve
reader comprehension without weakening editorial integrity. An operational
dashboard should support timely, correct response. A scientific visualization
must preserve domain-specific meaning. There is no honest way to replace those
different outcomes with one generic quality score.

On a six-level maturity scale, the August 2026 field is mostly at **level 3,
usable in bounded contexts**, for data grounding, ordinary visual construction,
chart interpretation, and mixed-initiative revision. It is mostly at **level 2,
repeatable on bounded tests**, for integrity critique and delivered interaction.
It remains at **level 1, demonstrated but not established**, for reader or
decision improvement and for production efficiency, governance, and
maintenance as one complete episode.

That is substantial progress, but it is not close to the ideal on the dimensions
that determine whether a visualization should be trusted, shipped, or credited
with helping a person. The frontier has broadened faster than outcome evidence
has matured.

## What “perfect” means

Perfect does not mean maximum autonomy. In consequential environments, a
perfect system may refuse, expose uncertainty, request a definition, or preserve
a human publication decision. The target is **appropriate independence plus
appropriate control**, not removal of people from the process.

This definition borrows three established visualization frameworks rather than
inventing a new theory of success:

- [Munzner's nested model](https://cs.ubc.ca/labs/imager/tr/2009/NestedModel/NestedModel.pdf)
  separates domain problem and data, task and data abstraction, visual encoding
  and interaction, and algorithm. Its central warning is that an upstream error
  cascades: perfect rendering cannot rescue the wrong question or abstraction.
- [Brehmer and Munzner's task
  typology](https://www.cs.ubc.ca/labs/imager/tr/2013/MultiLevelTaskTypology/)
  separates why the work is undertaken, what data and outputs it acts on, and
  how the visualization supports it.
- [Lam and colleagues' seven evaluation
  scenarios](https://petra.isenberg.cc/wiki/pmwiki.php?n=MyUniversity.EvaluationScenarios)
  distinguish work-practice, analysis and reasoning, communication,
  collaboration, user performance, user experience, and algorithm evaluation.
  The evaluation question should determine the method; a convenient metric
  should not determine the claim.

Together they imply a fixed common spine and context-specific validation. The
same artifact can be technically correct yet solve the wrong problem, impair a
workflow, fail its reader, or become impossible to maintain.

## The maturity scale

The score is the highest maturity supported by public evidence for a defined
job and population. It is not a normalized benchmark percentage.

| Level | Name | Evidence required |
| ---: | --- | --- |
| **0** | Not demonstrated | No credible public evidence that the system can perform the job. |
| **1** | Demonstrable | Selected examples or an early study show the capability can occur. Reliability and boundary conditions are not established. |
| **2** | Repeatable | Held-out tasks succeed under declared inputs, model, tools, and scoring. The environment remains bounded. |
| **3** | Usable | Intended practitioners can steer, inspect, correct, and accept the work at acceptable total cost in a declared context. |
| **4** | Deliverable | The work survives production data, interaction, responsiveness, accessibility, governance, handoff, and update requirements with low hidden repair. |
| **5** | Outcome-proven | Compared with the relevant baseline, the system improves the intended human outcome and appropriately calibrates uncertainty, refusal, and escalation. |

Level 5 is not required for every internal component. The common ideal is level
4 across the first seven and ninth dimensions, plus level 5 for the human
outcome. Autonomy is a deployment choice, not a seventh maturity level.

The evidence confidence is reported separately:

- **High:** multiple direct studies or one strong benchmark with clear limits;
- **Medium:** bounded experiments, small studies, or converging mechanism
  evidence; and
- **Low:** demonstrations, product contracts, anecdotes, or a persistent absence
  that has not been tested directly.

No scores are averaged. A mean would let strength in code generation conceal a
failure in integrity or reader outcome, and different contexts require different
weights.

## The nine dimensions

| Dimension | What level 4 requires | What level 5 adds |
| --- | --- | --- |
| **Task and audience framing** | The system identifies the real job, audience, stakes, ambiguity, and human authority before choosing an output. | The framing demonstrably improves the target outcome and refusal or escalation is calibrated. |
| **Data, semantics, and provenance** | Measures, joins, units, filters, uncertainty, permissions, freshness, and source lineage remain correct and inspectable through delivery. | Better decisions or understanding can be attributed to this grounded treatment. |
| **Visual construction fidelity** | Data transformations, marks, encodings, labels, layout, and renderer behavior are correct across the supported design space. | The constructed result outperforms the relevant professional baseline for the human objective. |
| **Visual interpretation and reasoning** | The system reads ordinary and difficult charts, multi-chart figures, and relevant domain conventions with calibrated selective risk. | Human outcomes improve because the interpretation is used appropriately. |
| **Integrity critique and uncertainty** | Misleading encodings, unsupported claims, perceptual defects, uncertainty, and evidence limits are detected without damaging sound work. | The intervention measurably improves calibrated trust or prevents consequential error. |
| **Iterative steering and repair** | People can inspect state, make local changes, branch, revert, and repair defects without regression. | The interaction improves the human's reasoning or accepted result, not only convenience. |
| **Interaction, responsiveness, and accessibility** | Intended behavior survives browser state, desktop and phone layouts, keyboard and assistive paths, export, and authentication. | Representative users complete their real tasks better than under the relevant baseline. |
| **Reader and decision outcomes** | This is an intermediate requirement only: defined readers and decision-makers are actually observed in the delivered context. | Comprehension, retention, decision quality, calibrated trust, or learning improves without hidden harm. |
| **Production efficiency, governance, and maintenance** | Total human-plus-machine cost, review, security, deployment, refresh, ownership transfer, regression, and later correction are acceptable. | The lifecycle produces better sustained organizational or public outcomes than the alternative. |

## Current common-core scorecard

This is a scorecard of the strongest credible **publicly demonstrated field
maturity**, not the median product or ordinary user's experience. “3” means
usable in at least a bounded declared context; it does not mean reliable across
all use cases.

| Dimension | Aug. 2026 | Confidence | Why it is not higher |
| --- | ---: | --- | --- |
| Task and audience framing | **2 / 4** | Medium | Systems can infer analytical tasks and support clarification, but evidence that they identify the right human purpose, audience, consequence, and authority remains thin. Sources: [NL4DV-LLM](https://arxiv.org/abs/2408.13391); [Data Formulator 2](https://arxiv.org/abs/2408.16119). |
| Data, semantics, and provenance | **3 / 4** | Medium | Database pipelines, visible transformed tables, and governed semantic models support usable grounding. Authoritative source choice, correction state, and end-to-end provenance remain inconsistent. Sources: [nvAgent / VisEval](https://aclanthology.org/2025.acl-long.960/); [Data Formulator 2](https://arxiv.org/abs/2408.16119); [Looker’s documented semantic grounding](https://docs.cloud.google.com/looker/docs/conversational-analytics-overview) as a capability contract, not an outcome study. |
| Visual construction fidelity | **3 / 4** | High | Conventional static work is usable and constrained systems can be excellent. Real-data, multi-turn, scientific, and interactive benchmarks still show material semantic and visual failures. Sources: [RealChart2Code](https://arxiv.org/abs/2603.25804); [Raiven](https://arxiv.org/abs/2604.10008); [DashArena](https://arxiv.org/abs/2608.10567). |
| Visual interpretation and reasoning | **3 / 4** | High | Basic chart QA has improved substantially. On 100 deliberately hard professional tasks, the best tested configuration reached 45% mean pass@1; multilingual, document, and multi-chart gaps remain. Sources: [Chartography](https://arxiv.org/abs/2608.10677); [POLYCHARTQA](https://aclanthology.org/2026.acl-long.2043/); [Chart-MRAG](https://aclanthology.org/2026.acl-long.1164/); [Beyond Single Plots](https://aclanthology.org/2026.findings-acl.1764/). |
| Integrity critique and uncertainty | **2 / 4** | High | On misleading visualizations, average model accuracy was 26.4% against a 25.6% random baseline. Table conversion helped some cases and harmed ordinary charts when extraction failed. Sources: [Protecting MLLMs against misleading visualizations](https://aclanthology.org/2026.acl-long.377/); [Misviz](https://arxiv.org/abs/2508.21675); [VisJudge](https://arxiv.org/abs/2510.22373). |
| Iterative steering and repair | **3 / 4** | Medium | Visible state, direct controls, branches, atomic edits, and execution feedback are usable in bounded systems. Regressive editing and diminishing critic returns remain common. Sources: [Data Formulator 2](https://arxiv.org/abs/2408.16119); [NL2Dashboard](https://arxiv.org/abs/2601.06126); [DashChat](https://arxiv.org/abs/2504.12865); [RealChart2Code](https://arxiv.org/abs/2603.25804). |
| Interaction, responsiveness, and accessibility | **2 / 4** | High | Dashboard state and replay are now measurable, but no DashArena model exceeded 86% render success or 74% trajectory replay. Mobile, responsive, keyboard, and assistive use remain largely outside AI benchmarks. Sources: [DashboardQA](https://aclanthology.org/2026.findings-eacl.177/); [Dashboard2Code](https://arxiv.org/abs/2607.04727); [DashArena](https://arxiv.org/abs/2608.10567). |
| Reader and decision outcomes | **1 / 5** | High | Selected studies make benefit and harm measurable, but no captured study shows professional AI-assisted production improving representative reader or decision outcomes across contexts. Sources: [proactive versus passive assistance](https://arxiv.org/abs/2409.11645); [tactile charts plus LLM assistance](https://arxiv.org/abs/2607.23065). |
| Production efficiency, governance, and maintenance | **1 / 4** | Medium | Build diaries and interviews show draft-time gains and substantial correction or handoff costs. No captured study closes accepted delivery, total cost, governance, later updates, and ownership transfer together. Sources: [Vibe Visualizing](https://arxiv.org/abs/2606.08914); [AI-supported end-user development](https://iris.unibs.it/handle/11379/630805); [Data Formulator 2](https://arxiv.org/abs/2408.16119). |

The shortest honest summary is therefore:

```text
bounded creation and reading: usable
integrity and delivered interaction: repeatable
human outcome and operational afterlife: demonstrated, not established
```

## What the ideal requires in different contexts

The common dimensions stay fixed. The release gates and outcome evidence change.

| Context | What perfect looks like | Strongest current evidence | Binding gaps in 2026 |
| --- | --- | --- | --- |
| **Explain and present** | A source-grounded account survives editorial review, mobile and accessible delivery, and improves comprehension or calibrated trust for defined readers. | Static construction is level 3; selected reader-assistance and misinformation studies make outcomes measurable. | Integrity **2**, delivery **2**, reader outcome **1**. |
| **Explore and discover** | An analyst forms and tests valid questions, sees transformations and alternatives, can branch and reverse course, and reaches sound findings with lower total effort. | Grounding and mixed-initiative steering reach level 3 in bounded tools and studies. | Critique **2**; discovery quality and total-effort outcome remain **1–2**. |
| **Monitor and respond** | Governed, fresh measures and thresholds support timely detection and response; false alarms, misses, access, uptime, escalation, and handoff are measured. | Semantic grounding can reach level 3 inside governed products. | Stateful delivery **2**; operational outcome and maintenance **1**. |
| **Compare and decide** | Fair alternatives, definitions, uncertainty, and counterevidence improve consequential decisions while preserving an audit trail and accountable owner. | Data grounding and chart interpretation reach level 3 in bounded analytical environments. | Integrity **2**; decision outcome **1**. |
| **Produce, revise, and reuse** | An accepted artifact is created at lower total cost, edits remain local, tests pass in the delivered surface, and another person can update it later. | Construction and steering reach level 3; interaction is repeatable at level 2. | Delivered behavior **2**; lifecycle evidence **1**. |
| **Scientific and professional visualization** | Domain conventions, coordinate systems, uncertainty, and specialized transformations survive expert review and reproducible delivery. | A restricted scientific language can approach level 4 for fully specified reproduction; hard professional interpretation remains near level 2. | Domain framing, integrity, open-ended reasoning, and reader outcome. |
| **Reader assistance, learning, and accessibility** | Assistance adapts to the actual reader and device, improves immediate and delayed unassisted performance, and does not create dependence or false confidence. | A few controlled studies establish level 1–2 mechanisms and measurable outcomes. | Delivered accessibility, representative sampling, delayed transfer, and harm. |

These rows are not product ratings. They show why the same global capability can
be sufficient for one bounded draft and unacceptable for another environment.

## How the scorecard was reconstructed through time

The retrospective series uses the same 2026 rubric at every date. Each score
asks: **what maturity had public evidence reached by that moment?** It does not
run current models on old benchmarks, award later knowledge retroactively, or
pretend that different benchmark percentages share a common scale.

Dates were selected when the public evidence surface changed materially:

- **2017:** declarative interactive grammars made structured construction and
  interaction programmable;
- **2019:** learned visualization generation and formal design constraints made
  automated generation and recommendation explicit;
- **2020:** natural-language queries produced inspectable analytic
  specifications;
- **2022:** human-written, real-chart question answering made visual reasoning
  repeatable;
- **2023:** LLM pipelines broadened goal, code, critique, and repair, while
  general multimodal models began reading renders;
- **2024:** render–inspect–repair loops and mixed-initiative authoring added
  executable and human-steerable evidence;
- **2025:** coding agents, database visualization, dashboard-specific systems,
  and targeted feedback expanded usable workflows; and
- **August 2026:** professional, multilingual, multi-chart, document, interactive
  state, replay, and judge-reliability benchmarks broadened the frontier.

## Historical scorecards on the fixed scale

| Dimension | 2017 | 2019 | 2020 | 2022 | 2023 | 2024 | 2025 | Aug. 2026 | Ideal |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| Task and audience framing | 0 | 1 | 1 | 1 | 2 | 2 | 2 | **2** | 4 |
| Data, semantics, and provenance | 1 | 1 | 1 | 1 | 2 | 3 | 3 | **3** | 4 |
| Visual construction fidelity | 1 | 1 | 1 | 1 | 2 | 2 | 3 | **3** | 4 |
| Visual interpretation and reasoning | 0 | 0 | 1 | 2 | 2 | 2 | 3 | **3** | 4 |
| Integrity critique and uncertainty | 1 | 1 | 1 | 1 | 1 | 2 | 2 | **2** | 4 |
| Iterative steering and repair | 1 | 1 | 1 | 1 | 2 | 3 | 3 | **3** | 4 |
| Interaction, responsiveness, and accessibility | 1 | 1 | 1 | 1 | 1 | 1 | 2 | **2** | 4 |
| Reader and decision outcomes | 0 | 0 | 0 | 0 | 0 | 1 | 1 | **1** | 5 |
| Production efficiency, governance, and maintenance | 0 | 0 | 0 | 0 | 0 | 1 | 1 | **1** | 4 |

Three patterns matter more than any cell:

1. **The representation substrate arrived early.** [Vega-Lite](https://vis.mit.edu/pubs/vega-lite/)
   could compile concise interactive specifications in 2017. The missing
   capability was not rendering; it was reliable mapping from human purpose and
   real data into the specification.
2. **Early learned generation was demonstration-grade.** [Data2Vis](https://arxiv.org/abs/1804.03126)
   produced simple univariate and bivariate Vega-Lite charts, but reported
   failure conditions in roughly 15–20% of tests, phantom fields, and only a
   qualitative comparison. [Draco](https://dig.cmu.edu/publications/2018-draco.html)
   made design knowledge testable as constraints, a durable mechanism that did
   not by itself understand purpose or outcome.
3. **The scorecard moved when evidence crossed a maturity boundary.**
   [NL4DV](https://www.tableau.com/research/publications/nl4dv-toolkit-generating-analytic-specifications-data-visualization-natural)
   showed structured natural-language specification in 2020 but explicitly
   lacked a large labeled benchmark. [ChartQA](https://aclanthology.org/2022.findings-acl.177/)
   made real-chart reasoning repeatable in 2022. [LIDA](https://aclanthology.org/2023.acl-demo.11/)
   made grammar-agnostic LLM generation repeatable on 57 datasets, but its
   quality evaluator read code rather than the rendered chart and was not
   calibrated to human outcomes.

The apparent plateau from 2025 to 2026 is not a claim that models stopped
improving. The 2026 frontier expanded sideways into much harder objects and
contexts. An ordinal maturity score changes only when evidence crosses a
threshold; large within-level gains and harder new benchmarks belong in the
supporting evidence, not in invented decimals.

## Forecasts translated into score movements

The existing forecasts resolve against explicit future evidence. Translating
them into this scorecard clarifies what each success would—and would not—change.

| Resolution test | Existing forecast | Score movement if the test passes | What would remain unchanged |
| --- | ---: | --- | --- |
| By Aug. 2027, exceed 70% accepted success on at least 500 real-data static-chart tasks with executable and human-calibrated visual checks. The closest current evidence includes [RealChart2Code](https://arxiv.org/abs/2603.25804), [Text2Vis](https://aclanthology.org/2025.emnlp-main.1622/), and [Raiven](https://arxiv.org/abs/2604.10008). | 78% probability / 56% confidence | Visual construction **3 → 4** for the declared static-chart population. | Reader outcome, maintenance, and broader interactive delivery. |
| By Aug. 2027, exceed 60% on [DashboardQA](https://aclanthology.org/2026.findings-eacl.177/) or a harder successor through executed replayable interaction. | 64% / 52% | Interaction and delivery **2 → 3** for bounded dashboard tasks. | Mobile, accessibility, production uptime, and real decision quality. |
| By Aug. 2027, an independent critic or chart-tool layer adds at least ten points to defect detection or successful repair without lowering total acceptance. Current anchors: [Protecting MLLMs against misleading visualizations](https://aclanthology.org/2026.acl-long.377/), [VisJudge](https://arxiv.org/abs/2510.22373), and [ChartAgent](https://aclanthology.org/2026.acl-long.843/). | 72% / 59% | Integrity critique **2 → 3**; confidence in repair at level 3 increases. | Human outcome unless the repaired artifacts are tested with people. |
| By Aug. 2027, an environment-specific package adds ten points on an independently verified non-scientific task set. The bounded visualization precedent is [SciVisAgentSkills](https://arxiv.org/abs/2606.05525). | 67% / 55% | No automatic core-score increase. It raises confidence that level 3 can transfer to another declared context. | Global framing, delivery, and outcomes; a technique win is not a maturity win by itself. |
| By Aug. 2028, two environments show at least 20% less total human time to an accepted artifact without worse correctness, reader outcome, or later update performance. | 43% / 42% | Production efficiency and maintenance **1 → 3** in those environments. | Broad governance and contexts not included in the studies. |
| By Aug. 2029, AI-assisted professional production improves representative-reader comprehension or calibrated trust by at least five points in two consequential contexts. Current bounded precedents: [proactive assistance](https://arxiv.org/abs/2409.11645) and [tactile charts plus LLM assistance](https://arxiv.org/abs/2607.23065). | 34% / 37% | Reader outcome reaches **5** in the tested contexts; the conservative cross-context score moves **1 → 3**. | Untested audiences, decisions, accessibility, and long-term retention. |
| By Aug. 2029, a system publishes autonomously across three environments while meeting source, interaction, mobile, accessibility, reader, and update gates without human acceptance. | 14% / 34% | Delivery **2 → 4** and production/governance/maintenance **1 → 4** for the tested environments. | Autonomy does not prove that other framing or outcome contexts are solved. |

If all four 2027 tests pass, the visible frontier would change from “usable
creation, repeatable critique and interaction” to “deliverable static creation,
usable critique, and usable bounded interaction.” It would still leave the two
largest gaps—human outcomes and operational afterlife—essentially untouched.

The forecast events are correlated, so their probabilities should not be added
or multiplied into one expected score. The scorecard should move only when the
named resolution evidence exists.

## How to maintain the scorecard

Every future update should preserve five records:

1. the fixed maturity rubric and ideal target;
2. the exact population, task, environment, model, harness, and tools;
3. the evidence that justifies a threshold crossing;
4. the contexts to which the score does not transfer; and
5. the prior scorecard as a dated historical snapshot.

A benchmark improvement may change a score, confidence, both, or neither. A
larger test that confirms the same maturity should raise confidence without
inventing a higher level. A new hard distribution may reveal a narrower context
boundary without implying that all earlier capabilities regressed. A field study
may raise outcomes while leaving underlying model scores unchanged.

The most valuable missing evidence is not another aggregate chart leaderboard.
It is a linked set of complete episodes: purpose and audience, authoritative
data, accepted artifact, delivered interaction, representative human outcome,
total effort, later correction, and ownership transfer.

## Source map

The links in the score rows above lead to the evidence closest to each
judgment. The groups below explain how those sources enter the scorecard. A
source supports only the named claim; its presence is not an endorsement of a
larger product, model, or field-wide conclusion.

| Evidence block | Primary sources | What they support here | What they do not support |
| --- | --- | --- | --- |
| **Definition of the ideal** | [Munzner’s nested model](https://cs.ubc.ca/labs/imager/tr/2009/NestedModel/NestedModel.pdf); [Brehmer and Munzner’s task typology](https://www.cs.ubc.ca/labs/imager/tr/2013/MultiLevelTaskTypology/); [Lam et al.’s evaluation scenarios](https://petra.isenberg.cc/wiki/pmwiki.php?n=MyUniversity.EvaluationScenarios) | Separate purpose, abstraction, representation, interaction, algorithm, work practice, user performance, and human outcome before choosing an evaluation. | The nine dimensions or 0–5 maturity assignments; those are this report’s synthesis. |
| **2017–2020 historical anchors** | [Vega-Lite](https://vis.mit.edu/pubs/vega-lite/); [Data2Vis](https://arxiv.org/abs/1804.03126); [Draco](https://dig.cmu.edu/publications/2018-draco.html); [NL4DV](https://www.tableau.com/research/publications/nl4dv-toolkit-generating-analytic-specifications-data-visualization-natural) | Declarative interaction, learned chart generation, constraint-based recommendation, and natural-language analytic specifications had become demonstrable and inspectable. | Reliable framing, real-world delivery, or reader benefit. |
| **2022–2023 historical anchors** | [ChartQA](https://aclanthology.org/2022.findings-acl.177/); [LIDA](https://aclanthology.org/2023.acl-demo.11/) | Real-chart question answering and multi-stage LLM visualization generation crossed into repeatable public evaluation. | Professional chart-reading reliability, rendered-chart human quality, or downstream outcomes. |
| **Current construction, grounding, and steering** | [Data Formulator 2](https://arxiv.org/abs/2408.16119); [nvAgent / VisEval](https://aclanthology.org/2025.acl-long.960/); [RealChart2Code](https://arxiv.org/abs/2603.25804); [Raiven](https://arxiv.org/abs/2604.10008); [DashArena](https://arxiv.org/abs/2608.10567) | Visible transformed data, structured database composition, harder real-data generation, constrained scientific compilation, and open-ended interactive dashboard attempts. | One transferable success rate across contexts, production maintenance, or reader outcomes. |
| **Current reading, integrity, and delivery** | [Chartography](https://arxiv.org/abs/2608.10677); [Protecting MLLMs against misleading visualizations](https://aclanthology.org/2026.acl-long.377/); [DashboardQA](https://aclanthology.org/2026.findings-eacl.177/); [Dashboard2Code](https://arxiv.org/abs/2607.04727); [POLYCHARTQA](https://aclanthology.org/2026.acl-long.2043/) | Difficult professional reading, misleading-chart vulnerability, dashboard navigation, hidden interaction state, and multilingual chart QA remain material bottlenecks. | The failure rate on ordinary charts or complete mobile, assistive, authenticated, and reader-tested delivery. |
| **Human outcome and lifecycle evidence** | [Proactive versus passive assistance](https://arxiv.org/abs/2409.11645); [Touching or Chatting](https://arxiv.org/abs/2607.23065); [Vibe Visualizing](https://arxiv.org/abs/2606.08914); [AI-supported end-user development](https://iris.unibs.it/handle/11379/630805) | Bounded comprehension, preference, accessibility, novice production, correction, and organizational experience can be studied directly. | Representative professional reader benefit or a complete production-to-maintenance episode. |

The comprehensive [benchmark
crosswalk](/reports/benchmark-crosswalk/)
records each benchmark’s input, output, data source, grader, human context,
supported claim, and important nonclaim. The [forecast
report](/reports/capability-forecast/)
owns the probabilities, confidence judgments, resolution tests, and annulment
conditions translated into score movements above.

## Judgment note

The scores are editorial synthesis constrained by those sources. They are not
the output of a validated psychometric instrument, and they should not be used
to rank products or models. Their purpose is to make the definition of progress,
the remaining distance, and the evidence required to move explicit.

## Update log

- **2026-08-14 — Initial public edition.** Defined the ideal-state dimensions,
  fixed maturity rubric, August 2026 field assessment, historical
  reconstruction, contextual endpoints, and forecast score movements.
