Forecasts as score movements

The likely near-term gains improve bounded construction, critique, and interaction—not the two largest gaps.

A forecast moves a score only if its named resolution test passes. Probabilities are not combined because the events are correlated.

ByResolution evidenceIf it passesForecast

More than 70% accepted success on 500 real-data static-chart tasks with executable and human-calibrated checks.Current anchors: RealChart2Code ↗Text2Vis ↗Raiven ↗

Construction3 → 4

78%56% confidence

More than 60% on DashboardQA or a harder successor through executed replayable interaction.Current anchor: DashboardQA ↗

Delivery2 → 3

64%52% confidence

An environment-specific package adds ten independently verified points outside scientific visualization.Bounded precedent: SciVisAgentSkills ↗

Transfer evidenceConfidence ↑

67%55% confidence

Two environments show 20% less total human time to an accepted artifact without worse correctness, reader outcome, or later update.

Lifecycle1 → 3

43%42% confidence

Professional AI assistance improves representative-reader comprehension or calibrated trust by five points in two consequential contexts.Bounded precedents: Proactive assistance ↗Tactile + LLM study ↗

Outcome1 → 3 globaltested contexts reach 5

34%37% confidence

Autonomous publication clears source, interaction, mobile, accessibility, reader, and update gates in three environments.

Delivery · lifecycle2 → 4 · 1 → 4

14%34% confidence