StudyWhat it measuredWhat it foundWhat it cannot establish
Analyst verificationCHI 2024 · 22 analysts · 52 workflows
Use of explanation, code, original and intermediate data, results, and visual summaries
People began with procedure, then often moved to data after noticing trouble; expertise shaped the artifact they trusted
Prepared tasks at one company; data transformations rather than complete visualization projects
Data Formulator 2CHI 2025 · eight participants · 16 charts each
Mixed direct manipulation and natural language on reproduction tasks
All completed; visible transformed data, code, explanations, history, and branches supported different verification styles
Six needed hints; no baseline, open exploration, self-owned data, or long-term use
Dashboard prototyping2025 preprint · 10 formative + 28 evaluation participants
Rapid mockup generation, structured edits, and comparison with a lightly taught Tableau condition
Participants valued speed, simulated data, history, and a concrete object for negotiation
Pre-data prototypes, not analytical correctness, deployment, or ongoing dashboard use
Visualization adviceTOCHI 2025 · 119 forum questions + 12 practitioners
AI versus human-expert design advice, ratings, interviews, and practitioner preference
AI helped enumerate ideas; practitioners preferred experts for accuracy, context, adaptability, and actionable advice
Practitioner sessions used 2023-era GPT-3.5; raw capability comparison is dated
Visualizationary2024 preprint · 13 designers + three expert raters
At least five versions of a self-selected visualization over a three-to-five-day window with LLM and perceptual critique
Final work improved 3.69/5 on average; intermediate and expert designers converted feedback into useful edits more readily
Roughly 90–150 minutes of observed work each; no critique baseline, production delivery, or later maintenance
Steering and verificationUIST 2024 · 18 analysts · 108 task episodes
Correction, premature acceptance, non-completion, task time, hints, and perceived control across conversational and decomposed interfaces
Seven episodes were not completed and 31 were declared complete with an issue remaining; structure improved perceived control, not detected success or time
Tasks were engineered to contain model errors and stopped at 15 minutes; no publication, maintenance, or field abandonment