Measured evidence

Speed, completion, confidence, and correctness can move in different directions.

The studies below answer different questions. Their denominators and limits remain attached; the numbers should not be pooled into one score.

Randomized public-health exercise · 2025 · 30 analyzed participants

The integrated tool was faster. Its work was less often free of serious errors.

The exercise compared integrated ChatGPT analysis with an R/Stata-plus-ChatGPT workflow on simulated epidemiological data. Overall scores were not significantly different.

Median completion timelower is better
Integrated AI38 min
Distributed tools45 min
Submissions free of serious errorshigher is better
Integrated AI6.7%
Distributed tools26%
Source ↗ Small, underpowered, 45-minute study with simulated data and no non-AI control. The bar scales are separate and labeled.
Vibe Visualizing · 2026 preprint

Novices could make charts. They could not reliably tell when the work had failed.

Twenty visualization novices completed 60 ChatGPT sessions producing 175 charts. The study observed prompting, chart quality, interpretation, verification, and repair—not just whether an image appeared.

Fatal task noncompliance52 of 60 sessions

The output omitted or contradicted a required element.

Incorrect insight recorded12 of 60 sessions

Poor chart design contributed to most incorrect-insight cases discussed.

Unusable chart22 of 175 charts

Every chart had at least one coded design flaw.

Meanwhile mean confidence was 3.73/5 and satisfaction 3.93/5. Only three verification attempts were observed.

Newer models produced fewer design flaws on replayed initial prompts, but richer interfaces added latency, broken controls, blank renders, and unverifiable interpretations. Better generation changed the failure surface; it did not remove the verification problem.

Source ↗ Small controlled datasets; the Gemini and Claude comparison replayed prompts and was not a live user study.
StudyWhat it measuredWhat it foundWhat it cannot establish

Analyst verificationCHI 2024 · 22 analysts · 52 workflows

Use of explanation, code, original and intermediate data, results, and visual summaries

People began with procedure, then often moved to data after noticing trouble; expertise shaped the artifact they trusted

Prepared tasks at one company; data transformations rather than complete visualization projects

Data Formulator 2CHI 2025 · eight participants · 16 charts each

Mixed direct manipulation and natural language on reproduction tasks

All completed; visible transformed data, code, explanations, history, and branches supported different verification styles

Six needed hints; no baseline, open exploration, self-owned data, or long-term use

Dashboard prototyping2025 preprint · 10 formative + 28 evaluation participants

Rapid mockup generation, structured edits, and comparison with a lightly taught Tableau condition

Participants valued speed, simulated data, history, and a concrete object for negotiation

Pre-data prototypes, not analytical correctness, deployment, or ongoing dashboard use

Visualization adviceTOCHI 2025 · 119 forum questions + 12 practitioners

AI versus human-expert design advice, ratings, interviews, and practitioner preference

AI helped enumerate ideas; practitioners preferred experts for accuracy, context, adaptability, and actionable advice

Practitioner sessions used 2023-era GPT-3.5; raw capability comparison is dated

Visualizationary2024 preprint · 13 designers + three expert raters

At least five versions of a self-selected visualization over a three-to-five-day window with LLM and perceptual critique

Final work improved 3.69/5 on average; intermediate and expert designers converted feedback into useful edits more readily

Roughly 90–150 minutes of observed work each; no critique baseline, production delivery, or later maintenance

Steering and verificationUIST 2024 · 18 analysts · 108 task episodes

Correction, premature acceptance, non-completion, task time, hints, and perceived control across conversational and decomposed interfaces

Seven episodes were not completed and 31 were declared complete with an issue remaining; structure improved perceived control, not detected success or time

Tasks were engineered to contain model errors and stopped at 15 minutes; no publication, maintenance, or field abandonment

Analyst verification ↗ · Data Formulator 2 ↗ · DashChat ↗ · Visualization advice ↗ · Visualizationary ↗ · Steering and verification ↗