What would unlock the next capabilities in AI-assisted data visualization?
Status: forecast from an evidence cut dated 14 August 2026. The probabilities are forecasts, not research findings. Each is paired with a dated resolution test so that the argument can be corrected rather than quietly rewritten.
Companions: The state of AI-assisted data visualization research, The practitioner and reader experience, Human skills and banked gains, and Specialized vision models for data visualization.
Executive assessment
The next step is not “better charts from prompts.” That capability is already available for conventional, low-stakes drafts. The remaining work is to make an AI-assisted result grounded, inspectable, correctable, deliverable, and effective for its actual reader.
Those properties unlock in sequence because each depends on evidence that the previous one does not supply:
- Grounding: Does the system understand the question, definitions, data, audience, and intended use?
- Inspectable construction: Can a person see and edit the transformation, chart state, assumptions, and alternatives?
- Verification: Do source, specification, render, interaction, mobile, and accessibility checks agree?
- Repair: Can the system correct a local defect without damaging something that was already right?
- Delivery: Does the artifact survive its browser, export, governance, handoff, and later update?
- Reader outcome: Does the intended person understand, decide, learn, or act better—and with appropriately calibrated trust?
Raw model improvement matters at every step, particularly for fine visual grounding, long-horizon consistency, interaction planning, and recovery from tool failure. But much of the strongest progress since 2023 did not come from a larger model alone. It came from changing what the model could see and do: converting a chart into an inspectable table, executing calculations, rendering and looking at the output, exposing transformed data, constraining generation through a language and compiler, attaching chart-specific tools, or supplying a small package of applicable procedural knowledge.
That history supports a near-term forecast with two speeds:
- Fast progress in bounded environments. Static chart work, narrow visual criticism, dashboard interaction, and specialist adapters already have measurable gaps and executable feedback. They are likely to improve over the next 12–24 months through a combination of stronger models and better technique.
- Slow progress in situated outcomes. Production productivity, editorial judgment, representative-reader benefit, governance, and autonomous publication lack both reliable machine signals and adequate field studies. Better models help, but they do not reveal facts, authority, consequences, or human reactions that the system never observes.
The most plausible near-future breakthrough is therefore not one universal visualization agent. It is a more selective system: a strong general model working over source data, semantic state, and rendered output; invoking small specialists and deterministic checks at known bottlenecks; and escalating when the evidence needed to decide is unavailable.
What “unlock” means
A capability is not unlocked when a model produces an impressive example. It is unlocked when a defined population can complete a defined job under the conditions that matter for that job.
| Level | What would count | What would not count |
|---|---|---|
| Demonstrable | a system can produce or interpret selected examples | a polished screenshot, cherry-picked dialogue, or vendor demonstration |
| Repeatable | held-out tasks succeed under a declared model, harness, cost, and evaluation contract | an aggregate score with no failure distribution or artifacts |
| Usable | intended practitioners can steer, inspect, correct, and accept the work at lower total cost | time to first draft or self-reported ease alone |
| Deliverable | the artifact survives data, browser, mobile, accessibility, handoff, governance, and update requirements | code execution, one desktop render, or creator satisfaction |
| Outcome-improving | readers, analysts, or decision-makers do better than under the relevant human-only baseline | model chart QA, an expert-style rating, or synthetic-reader approval |
| Autonomous | the system completes all applicable gates without hidden human repair or acceptance | a workflow that silently relies on benchmark gold answers, manual cleanup, or a human publication decision |
The field is already well into the first two levels for ordinary static charts. It has partial evidence at the usable level in bounded tasks. It has very little evidence at the deliverable and outcome-improving levels, and no captured system meets the autonomous definition.
The capability gates
The gates below are not stages every job must traverse in the same interface. They are distinct failure boundaries. A low-stakes exploratory chart may need a light delivery gate. A public health graphic may require every gate and several human owners.
| Gate | What fails now | The gain needed | What technique can contribute | What probably requires stronger underlying capability |
|---|---|---|---|---|
| Intent and semantic grounding | plausible charts answer a nearby question, use the wrong measure, or silently invent a denominator | preserve the requested question, local definitions, authoritative source, uncertainty, audience, and decision through every transformation | an explicit use contract; governed semantic models; visible assumptions; source and field selection; ambiguity escalation | better disambiguation and domain reasoning help, but an unobserved local definition cannot be inferred reliably from scale |
| Transformation correctness | aggregation, filtering, joins, missingness, and units disappear behind fluent code or prose | every plotted value can be traced to a source and inspectable operation | semantic chart specifications; visible intermediate tables; executable invariants; comparison with source totals | stronger coding and data reasoning reduce errors, especially on unfamiliar schemas, but cannot replace unavailable provenance |
| Fine visual grounding | models miss tiny labels, overlaps, occluded marks, color relations, flow direction, or the state of an interaction | reliably locate and describe the relevant mark, text, scale, state, and relationship at the needed resolution | crop, zoom, OCR, segmentation, chart-specific tools, multiple representations, and targeted specialist routing | better high-resolution perception, spatial reasoning, video or interaction understanding, and calibrated uncertainty |
| Context-preserving revision | a requested fix introduces a new error elsewhere; later turns lose constraints | local edits preserve correct data, layout, interaction, and prior decisions | versioned semantic state; diffs; regression checks; bounded edit operations; rollback; test-before-accept | longer reliable working state, stronger causal editing, and better recovery from partial failure |
| Critique and verification | a critic praises a plausible image while missing data or semantic defects, or produces generic advice that costs more to filter than to fix | detect the defects that matter in context, point to evidence, propose a repair, and know when it cannot judge | separate deterministic, perceptual, semantic, accessibility, and reader checks; specialist critics; uncertainty-triggered escalation | better cross-modal grounding and judgment; no model-only gain supplies source truth or representative human response |
| Interaction and environment control | agents mis-ground controls, take poor action sequences, fail to replay, or stop in the wrong state | reliably perceive state, plan actions, recover, and prove the resulting state | browser instrumentation; semantic accessibility trees; action constraints; replay; state assertions; failure-specific tools | stronger visual control, long-horizon planning, memory, and recovery under changing interfaces |
| Delivery and maintenance | a good draft fails on mobile, export, assistive technology, permissions, refresh, handoff, or later updates | accepted artifacts survive actual delivery and at least one change cycle | delivery-specific tests; responsive variants; provenance; governed APIs; update fixtures; explicit ownership | broader and more reliable computer use helps; organizational authority and changing external requirements remain environmental |
| Reader and decision outcome | creator success and model judgment are used as substitutes for understanding, calibrated trust, or action | demonstrate benefit with the intended readers in the intended setting | proactive scaffolding, alternative representations, accessible descriptions, instrumentation, and real reader tests | richer audience models may improve hypotheses; representative response and consequential acceptance remain empirical human evidence |
Two implications follow.
First, a stronger model can move several gates at once, but it cannot collapse them into one score. Second, the nearer a problem is to observable state and an executable test, the more likely a technique can expose capability that already exists. The nearer it is to local meaning, authority, or human response, the less likely prompt or model improvements alone are to settle it.
What this unlocks in different environments
Exploratory analysis
The next useful capability is not autonomous discovery. It is an inspectable branching partner that can propose transformations and views, retain the analyst’s question and history, and make every candidate cheap to reject.
The critical gains are semantic grounding, transparent transformations, branching without state loss, and local correction. Stronger code generation helps, but mixed-initiative controls and visible data are at least as important. Data Formulator 2 demonstrates that this is a workable interaction pattern, while its eight-person reproduction study does not establish open-ended analytical discovery or a productivity advantage over direct tools.
Operational BI and dashboards
The next capability is a governed dashboard collaborator that operates over declared measures and permissions, can inspect and exercise interactions, and can prove the state it leaves behind.
The remaining gap is not just generation. DashboardQA contains 292 tasks on 112 interactive dashboards. Its best reported agent reached 38.69% accuracy. The failures concentrate in grounding interface elements, planning interaction trajectories, and reasoning over the resulting state. Better visual control is needed, but a semantic layer, browser replay, and state assertions can convert some of that open-ended problem into testable operations.
Editorial and journalistic graphics
The next capability is a source-grounded production collaborator, not an autonomous storyteller. It should preserve the reporting claim and uncertainty, generate and compare alternatives, produce responsive states, and make factual and visual review cheaper without manufacturing a narrative before the evidence supports it.
Raw visual and coding improvements will continue to reduce implementation cost. The hard unlocks are source custody, local editorial judgment, mobile reading, and representative-reader outcomes. None is available from the chart image alone. A useful system must know which questions require a reporter, editor, designer, accessibility specialist, or reader study rather than simulating all of them inside one model.
Scientific and geospatial visualization
The next capability is a specialist operating partner that knows the local file formats, arrays, projections, APIs, rendering state, and domain conventions—and can expose its operations for correction.
This is the environment with the clearest evidence for technique-led gains. Raiven constrains authoring through a restricted language and deterministic compiler. The SciVisAgentSkills study reports improvement across all ten tested suite–agent quality comparisons after adding tool-specific procedural packages. Those findings are bounded and mostly same-team, but they identify the mechanism: fragmented, versioned knowledge and executable operations are difficult to infer repeatedly and valuable to supply directly.
Learning and reader assistance
The next capability is an adaptive guide that teaches without hiding the evidence or replacing independent judgment.
In a randomized 117-person experiment, proactive questioning produced better post-intervention comprehension than a passive conversational agent or a data story. This is evidence for a particular scaffolding technique, not for a general synthetic reader or automatic explanation system. The necessary next gain is delayed transfer to unfamiliar charts, paired with calibration: people should become better at recognizing what they do and do not understand after the assistance is removed.
What improved because models improved—and what improved because technique improved
The distinction is not perfectly clean. Training a chart specialist changes weights; attaching a tool changes inference; a hosted coding harness may bundle planning, execution, and visual inspection around a stronger base model. The useful question is causal: would the gain still have appeared if the base checkpoint had stayed fixed?
Gains that are primarily model- or training-driven
- Image input and general visual reasoning became baseline capabilities. GPT-4V in 2023 made it practical for a general model to inspect a render. Later native multimodal models improved text recovery, layout fidelity, and broad chart reading enough that render–inspect–revise loops became ordinary harness operations rather than specialist research prototypes.
- Conventional static chart reproduction became much more faithful. The same-benchmark comparisons summarized in the field review show large movement in text and rendered-chart fidelity from 2023-era to 2025–2026 models, while execution moved much less. This is a real baseline shift, but it is not a clean annual learning curve.
- Broad scientific visualization reading is now plausible for some frontier models. A July 2026 study tested six models on 49 items spanning 18 scientific visualizations. Gemini 3.1 Pro Preview reached 88.6% overall, above the reported 75.6% human mean; GPT-5.4 and Claude Opus 4.6 were near that mean. Performance remained uneven on fine quantitative estimation, flow direction, and unfamiliar encodings.
- Data-centric chart training can compete with frontier scale. CharTide separates visual perception, code logic, and modality fusion during tuning and uses answer-invariance checks as a reinforcement signal. Its compact models surpassed GPT-4o and were competitive with GPT-5 on the selected chart-to-code benchmarks. This is a training-technique gain embodied in a model, not evidence that scale is irrelevant.
These gains expanded the set of feasible systems. They did not establish reliable real-data generation, non-regressive editing, interactive control, or reader benefit. On RealChart2Code, top models that averaged 91–96% on simpler chart-to-code benchmarks fell to roughly 51% on real data, complex multi-panel structures, and refinement tasks. The paper also identifies regressive editing: models repair the requested defect while damaging previously correct code or output.
Gains that are primarily technique-driven
The following results hold the base model fixed or compare matched conditions. Their metrics are not interchangeable, but together they show how much capability can be latent in a model and unavailable under a direct prompt.
| Date | Intervention | Matched result | What the technique added | Boundary |
|---|---|---|---|---|
| 2023 | DePlot converts a chart image into a table before language reasoning | 67.6% on ChartQA’s human-written split versus 38.2% for the prior end-to-end system | an inspectable numeric representation and a division of perception from reasoning | the table discards color, direction, geometry, and other visual evidence; later hard benchmarks exposed transfer limits |
| 2024 | MatPlotAgent adds planning, execution, debugging, and rendered visual feedback | GPT-4 moved from 48.86 to 61.16 on MatPlotBench; removing visual feedback reduced the agent result to 53.44; zero-shot chain-of-thought fell to 45.42 | a render–perceive–repair loop with genuinely new sensory evidence | author-built 100-task benchmark; GPT-4V judge; some base models regressed |
| 2024 | TinyChart executes Python for calculative questions | calculative accuracy moved from 56.64 with direct answers to 78.98 with program-of-thought | executable arithmetic instead of hidden visual calculation | older chart distributions; a simple router still left a gap to oracle routing |
| 2025 | Text2Vis adds targeted answer-and-code feedback | GPT-4o’s final pass rate moved from 26% to 42%; three examples alone left it at 26%, and retrieval plus examples reached 31% | feedback tied to answer correctness and executable code | adding visual feedback improved visual subscores but slightly lowered final pass to 41%; more feedback was not always better |
| 2026 | ChartAgent adds more than 40 chart-specific perception and calculation tools | the same GPT-4o orchestrator moved from 54.53% to 71.39% on ChartBench; chart-specific tools beat generic image tools by 30 points overall | active perception, localization, and executable calculation | half of 30 audited trajectories needed recovery or failed; single-chart QA is not authoring or publication |
| 2026 | VisDeception reconstructs structured chart metadata before answering | deception error fell .099→.024 for Gemini 2.5 Pro and .502→.255 for GPT-4o | an explicit semantic representation alongside the image | Gemini 2.5 Flash worsened .109→.135; a bad intermediate representation propagates error |
| 2026 | SkillsBench supplies curated procedural packages | mean pass rate across 18 model–harness configurations moved 33.9%→50.5% | applicable procedures, examples, scripts, and domain conventions | self-generated packages fell below baseline; comprehensive packages added only 0.7 points; the benchmark is broader than visualization |
The common pattern is not “agentic beats direct.” It is specific missing evidence or action beats a prompt that does not contain it. The techniques make values legible, calculations executable, output visible, state editable, or environment knowledge available. When an intervention adds only more prose, more turns, or an unreliable intermediate, its effect is weak or negative.
What the last 36 months do—and do not—let us project
There is no defensible visualization capability-doubling rate. Benchmarks became harder as models improved: templated single charts gave way to real data, scientific figures, deceptive designs, multi-view dashboards, interactions, and iterative repair. Later systems also bundle more tools and test-time work. A score from 2023 and one from 2026 often measure different objects.
The history nevertheless supplies three useful directional priors.
- When an adjacent capability has already crossed into the base model, productization follows quickly. Once models could see images, write useful code, and operate tools, render inspection and browser loops moved from bespoke research into ordinary agent harnesses over roughly 12–24 months.
- When the missing signal can be computed or exposed, technique can produce a discontinuous gain before the next model generation. Program execution, chart tools, structured data, source checks, and deterministic compilers all fit this pattern.
- When success depends on unobserved context or human response, benchmark progress does not imply a fast unlock. The literature still lacks representative-reader, accepted-delivery, maintenance, and consequential decision evidence. No rate can be estimated from zero qualifying studies.
The resulting forecast is intentionally asymmetric: relatively high probabilities for bounded static work, critics, interactions, and narrow adapters; lower probabilities for production productivity, reader benefit, and autonomy.
Dated forecasts
Probability is the chance that the stated resolution test will pass by its date. Confidence is how much evidence supports that probability. A high probability with moderate confidence means the direction looks likely but the measurement surface is still young.
| Resolution date | Forecast | Probability | Confidence | Why this level |
|---|---|---|---|---|
| Aug. 2027 | A general-model system exceeds 70% accepted-task success on at least 500 held-out, real-data, single-static-chart tasks with executable and human-calibrated visual checks. | 78% | 56% | conventional static generation is already strong; the hard benchmark sits near 51%, and both baseline and verification technique are moving |
| Aug. 2027 | An agent exceeds 60% on DashboardQA or a harder successor through executed, replayable interactions. | 64% | 52% | 38.69% leaves a large but concrete gap; GUI grounding and planning are improving outside visualization, and the task now has a public test |
| Aug. 2027 | An independently evaluated critic or chart-tool layer adds at least ten points to a current general model’s defect detection or successful repair without lowering total acceptance. | 72% | 59% | several same-model interventions already clear this magnitude; independence, repair, and non-regression are the missing conditions |
| Aug. 2027 | A compact environment-specific package adds at least ten points on an independently verified non-scientific visualization task set. | 67% | 55% | scientific and broad skill studies demonstrate the mechanism; ordinary chart, dashboard, editorial, accessibility, and geospatial replications remain absent |
| Aug. 2028 | Field evidence in two environments shows at least 20% less total human time to an accepted artifact, without worse correctness, reader outcome, or later update performance. | 43% | 42% | plausible productivity gains exist, but current studies rarely include direct-work, accepted-delivery, and maintenance measures together |
| Aug. 2029 | AI-assisted production improves representative-reader comprehension or calibrated trust by at least five points over human-only professional production in two consequential contexts. | 34% | 37% | proactive scaffolding is promising; the required professional baseline and multi-context reader evidence do not yet exist |
| Aug. 2029 | A system publishes autonomously across three environments while meeting source, interaction, mobile, accessibility, reader, and update gates without human acceptance. | 14% | 34% | every component is improving, but no captured system clears all gates in one environment; several gates require facts or authority outside the model |
These are not forecasts that “AI will be capable” in the abstract. They are forecasts that a public evidence surface will meet a declared threshold. A private system may cross earlier; a benchmark may lag. That measurement delay is part of the uncertainty.
Where new technique breakthroughs are most likely
1. Better representation bridges
The next useful representation is unlikely to be one universal chart table or one universal dashboard language. Different checks need different evidence: source data, transformed table, semantic specification, code, rendered image, localized marks, interaction state, and delivered page.
A promising technique will retain these views and route questions between them. For example, numeric fidelity can be checked against source data; small mark overlap against pixels; a dashboard filter against interaction state; and an editorial claim against source and story context. Current models are already capable of reasoning over many of these views when they are supplied. The discovery problem is how to expose the smallest sufficient bundle without overloading or contradicting the model.
2. Verifiers that generate useful repair evidence
Many current checks return a score or a complaint. The next step is a verifier that identifies the violated condition, localizes the responsible state, proposes a bounded change, and reruns the relevant tests. This can unlock reliable revision sooner than a general leap in taste or judgment.
The important distinction is between detection and repair. A critic that finds a problem but causes regressions is not a quality loop. Research should measure defects found, defects fixed, new defects introduced, human filtering time, and stopping behavior separately.
3. Active visual perception
ChartAgent’s tool ablation suggests that a current multimodal model can answer harder visual questions when it can crop, segment, annotate, localize axes, and calculate. Similar techniques can support authoring critique: zoom into a legend, compare a mobile crop, inspect a hover state, or trace a mark to its value.
The likely breakthrough is not a model staring longer at the same screenshot. It is a policy for deciding where to look next, which specialist or tool to invoke, and when the observation is too uncertain to accept.
4. Conditional routing with a no-scaffold path
Skills and multi-stage systems show large wins and real regressions. This makes selection itself a core capability. A future system should be able to use no additional scaffold when the base model is already strong, load a compact specialist when local knowledge is missing, and fall back when the specialist does not match the version or task.
This is a fertile technique problem for current models. Realistic skill retrieval materially underperforms oracle selection, so there is substantial headroom without any change to the underlying model. Success requires classification, compatibility checks, observed effect, and an explicit no-specialist control—not a larger registry.
5. State-preserving edit operations
Regressive editing is a sign that free-form regeneration is the wrong action space for some revisions. Constraining changes to semantic objects, typed transforms, layout variables, or verified patches can make a model’s existing editing ability safer. Version history and regression fixtures should let the system compare “fixed the request” with “preserved everything else.”
This is likely to improve before models become universally reliable at long multi-turn edits, because the system can reduce the number of facts the model must preserve implicitly.
6. Adaptive reader scaffolding
The proactive-comprehension experiment suggests that asking a reader targeted questions can produce a different outcome from waiting for questions. The next technique opportunity is to adapt the scaffold to observed misunderstanding without answering for the reader, then measure delayed unassisted transfer.
This does not imply that a model can simulate the audience. The model may help construct and deliver a test; the actual reader remains the source of evidence about comprehension, trust, accessibility, and action.
Techniques likely to lose value as baselines improve
Some scaffolds mainly compensated for earlier model or harness limitations. They should be treated as replaceable experiments rather than permanent doctrine.
- generic chart-choice primers and broad style reminders;
- zero-shot chain-of-thought as a universal default;
- large prompts that restate normal planning and coding behavior;
- fixed multi-agent role play without distinct evidence, tools, or context;
- repeated visual self-review by the same model with no new representation;
- comprehensive skill packages loaded for every task;
- exact workflow steps that duplicate an improving harness;
- extra iterations without an acceptance or stopping signal.
The negative evidence is already material. Plot2Code found no clear general advantage for chain-of-thought or Plan-and-Solve. MatPlotAgent’s zero-shot chain-of-thought reduced GPT-4’s score. Text2Vis examples alone were null. SkillsBench found only a 0.7-point mean gain for comprehensive packages, while self-generated packages fell below no-skill baselines. SWE-Skills-Bench found 39 of 49 packages made no difference and three version-mismatched packages reduced performance by as much as ten points.
The durable replacement is thinner: declare the job and evidence contract, expose local state, load one applicable specialist, run the checks, and stop.
What stronger models cannot reveal by themselves
Several remaining capabilities are often described as reasoning problems even though the missing object is external information.
| Missing object | Why model improvement is insufficient | Required source of evidence |
|---|---|---|
| local measure definition | the same label can have different institutional meanings | governed semantic layer or accountable domain owner |
| authoritative data and provenance | plausibility does not identify the accepted source or correction state | source system, custody record, and freshness check |
| editorial purpose and consequence | a model can propose frames but does not own the claim or harm | reporter, editor, subject expert, and publication policy |
| reader comprehension and trust | an expert-like judgment is not a representative human response | appropriately sampled readers completing realistic tasks |
| accessibility in use | generated descriptions and markup do not establish task completion | browser, assistive technology, and disabled users where stakes warrant |
| operational acceptance | code and screenshots do not establish permissions, refresh, handoff, or maintenance | live environment, owner acceptance, and later update evidence |
Models will get better at asking for, organizing, and reasoning over these facts. That is valuable. They will not make the facts unnecessary.
Signposts that should move the forecast
Raise the near-term forecasts if
- a RealChart2Code-class benchmark moves above 65% under independently checked semantics and visual fidelity, rather than execution alone;
- DashboardQA improves by at least 15 points across two model families with replayable trajectories and fewer grounding failures;
- a visualization critic improves human-validated repairs, not just ratings or another model’s score;
- a compact specialist replicates outside its authors’ environment and retains its gain across a model upgrade;
- edit systems report fewer new defects per successful requested repair; or
- two studies measure total human plus machine time to accepted delivery rather than time to first output.
Lower the forecasts if
- hard-distribution performance stays flat while older chart benchmarks saturate;
- apparent agent gains disappear under equal compute, tool access, and model versions;
- realistic routing continues to erase most oracle-selected specialist gains;
- visual critics remain weak on dashboards or fail to localize repairable causes;
- regressive editing persists despite semantic state and regression tests; or
- field studies find that verification and correction cost consume draft-time savings.
Evidence that would change the long-range view
The long-range forecast should move substantially only when research crosses the missing outcome boundaries: representative readers, consequential decisions, accepted production delivery, maintenance, and delayed human learning. Another chart-generation leaderboard, by itself, should barely move the autonomous-publication probability.
Research program implied by the forecast
- Maintain a moving model-plus-harness baseline. Rerun the same tasks after meaningful model, computer-use, browser, or skill-system changes. Preserve prompts, tools, costs, outputs, and failures.
- Ablate one source of technique at a time. Compare direct baseline; explicit context contract; semantic state; deterministic checks; visual tools; one specialist; and full orchestration. Do not credit the full stack for a gain that one component produced.
- Measure repair, not only generation and detection. Inject realistic semantic, transform, visual, responsive, interaction, and accessibility failures. Count successful repair, new defects, abandonment, and human time.
- Cross capability profiles with environments. A domain expert with low coding skill, a visualization specialist new to the data, and a dashboard operator need different assistance and expose different failure modes.
- Follow artifacts through delivery and use. Include browser and mobile states, accessibility, handoff, one later update, and representative-reader tasks where the visualization is meant to communicate.
- Resolve forecasts in public evidence, not impression. Update the probability when a named signpost changes, record why it moved, and score it on the stated date.
Bottom line
The last three years show that models and techniques are complements, not rival explanations. Stronger models created the substrate: usable code, multimodal perception, long context, tool use, and broader visual literacy. Technique made those abilities dependable enough to matter on selected jobs by exposing state, supplying missing domain knowledge, executing operations, and closing the render–inspect–repair loop.
That pattern is likely to continue. Current and near-future models probably contain more usable visualization capability than a direct prompt reveals. The best technique discoveries will not be cleverer incantations. They will change the information and action structure around the model: plural representations, active perception, executable verification, bounded edits, conditional specialists, and explicit escalation.
This should unlock strong bounded collaborators relatively soon. It does not support a near-term forecast of autonomous, audience-aware publication. The distance between those outcomes is the distance between producing an artifact and proving that it is true, usable, maintainable, and helpful to the people for whom it exists.
Evidence and forecast note
The underlying research record contains 20 captured sources, seven verified synthesis claims, and seven dated forecast questions. Numeric comparisons are reported only within matched studies; no cross-benchmark score is treated as a learning curve. The forecasts distinguish probability from evidentiary confidence and include explicit annulment and resolution conditions.
The most load-bearing sources are:
- DashboardQA, 2026;
- RealChart2Code, 2026;
- Text2Vis, 2025;
- ChartAgent, 2026;
- MatPlotAgent, 2024;
- Raiven, 2026;
- Data Formulator 2, CHI 2025;
- Improving Steering and Verification in AI-Assisted Data Analysis, UIST 2024;
- SkillsBench, 2026;
- SWE-Skills-Bench, 2026; and
- The Effects of Generative AI Agents and Scaffolding on Visual Analytics Comprehension, 2024 preprint.
Update log
- 2026-08-14 — Initial public edition. Converted remaining capability gaps into measurable gates, separated model-driven from technique-driven gains, and recorded dated forecasts, signposts, falsifiers, and score movements.