Forecast from the August 2026 evidence cut

The next unlock is a chain of evidence, not a better first draft.

A chart capability becomes useful as it moves from grounded intent to inspectable construction, verification, repair, delivery, and reader outcome. A stronger model can move several links. It cannot replace evidence the system never sees.

  1. 01

    Ground

    Question, definitions, source, audience, stakes

  2. 02

    Construct

    Visible transforms, semantic state, alternatives

  3. 03

    Verify

    Source, values, render, interaction, delivery

  4. 04

    Repair

    Fix locally without introducing a regression

  5. 05

    Deliver

    Browser, mobile, accessibility, handoff, update

  6. 06

    Help

    Readers understand, decide, learn, or act better

Each link has a different acceptance test. Code execution cannot prove data fidelity. A clean render cannot prove interaction. Creator acceptance cannot prove reader comprehension.

Why the horizons differ

Observable, executable failures are likely to improve first.

Static chart generation, dashboard interaction, visual tools, and narrow adapters already have public tests and measurable gaps. Production productivity and reader benefit require field evidence that is mostly absent.

The percentages are dated probabilities that a declared public evidence test will pass—not estimates of how intelligent a future model will be. Confidence is lower where the field has no stable base rate.

Evidence thresholdByProbability

Reliable static work

At least 70% accepted success on 500+ real-data chart tasks with executable and human-calibrated visual checks.

78%56% confidence

Interactive dashboard reasoning

More than 60% on DashboardQA or a harder successor through executed, replayable interactions.

64%52% confidence

Critic or tool layer that repairs

An independent same-model test adds ten points to detection or repair without lowering total acceptance.

72%59% confidence

Production productivity

Two environments show 20% less total human time to an accepted, maintainable artifact without quality loss.

43%42% confidence

Reader benefit in consequential use

Two contexts improve representative-reader comprehension or calibrated trust over human-only professional production.

34%37% confidence

Autonomous publication

Three environments clear source, interaction, mobile, accessibility, reader, and update gates without human acceptance.

14%34% confidence