More than 70% accepted success on 500 real-data static-chart tasks with executable and human-calibrated checks.Current anchors: RealChart2Code ↗Text2Vis ↗Raiven ↗
78%56% confidence
Forecasts as score movements
A forecast moves a score only if its named resolution test passes. Probabilities are not combined because the events are correlated.
More than 70% accepted success on 500 real-data static-chart tasks with executable and human-calibrated checks.Current anchors: RealChart2Code ↗Text2Vis ↗Raiven ↗
78%56% confidence
More than 60% on DashboardQA or a harder successor through executed replayable interaction.Current anchor: DashboardQA ↗
64%52% confidence
An independent critic or chart-tool layer adds ten points to detection or repair without lowering total acceptance.Current anchors: Misleading-chart interventions ↗VisJudge ↗
72%59% confidence
An environment-specific package adds ten independently verified points outside scientific visualization.Bounded precedent: SciVisAgentSkills ↗
67%55% confidence
Two environments show 20% less total human time to an accepted artifact without worse correctness, reader outcome, or later update.
43%42% confidence
Professional AI assistance improves representative-reader comprehension or calibrated trust by five points in two consequential contexts.Bounded precedents: Proactive assistance ↗Tactile + LLM study ↗
34%37% confidence
Autonomous publication clears source, interaction, mobile, accessibility, reader, and update gates in three environments.
14%34% confidence