Technique area
Where evidence converges
Where design depends on context
What performs poorly or proves too little
What remains open
Intent and representation
Expose fields, transforms, encodings, interactions, and edits in an inspectable plan, language, contract, or native semantic model.
A compiler-backed DSL fits stable scientific grammars; direct generation may be proportionate for reversible one-off charts; governed BI can rely on maintained measures.
A formally valid plan can encode the wrong question. In Raiven, the DSL advantage was concentrated in scientific visualization rather than ordinary information graphics.
Which representation earns its cost for each environment, especially as direct model generation improves?
Verification and evaluation
Use computation for values, joins, filters, and runtime claims; evaluate the rendered artifact and, when interactive, its behavior and resulting state.
Static work may stop at final-context render inspection. Interactive work needs action coverage and replay. High-stakes work needs independent acceptance.
A nonblank render or clean execution is weak evidence: 21 of 30 cleanly executed DashArena failures still had semantic defects.
How should benchmark evidence connect to comprehension, retention, calibration, and real decisions?
Human steering
Visible transformed data, direct controls, history, branching, and reversion help people inspect and correct model work.
Exploration needs continuous steering; editorial explanation needs authorial control over claim and sequence; monitoring needs governed thresholds and response paths.
Synthetic personas are not reader evidence. A model-authored interaction trajectory describes intended use, not observed human behavior.
Do generated visuals help a defined audience understand, remember, or decide better over time?
Roles, feedback, and instructions
A separate role is defensible when it has different information, tools, or independent evidence: schema access, execution, visual inspection, or browser replay.
One capable model may be better for simple work; specialized roles may help when failure layers and evidence sources are genuinely distinct.
Repeated critics show diminishing returns. More instruction can reduce execution. Removing nvAgent’s processor slightly helped GPT-4o overall while hurting weaker and multi-table cases.
Equal-budget tests of role decomposition and behavioral with-and-without tests of skill packages remain rare.
Grounding and evidence claims
Source custody, semantic definitions, permissions, denominators, and data vintage must travel with the artifact.
The grounding source differs: governed semantic models in enterprise BI, explicit source packets in editorial work, and domain types in scientific systems.
Repository popularity, install count, saved API responses without renders, one vision-model score, or a validator that fails open cannot establish quality.
Public product evaluations stratified by semantic-model quality and realistic data conditions are still missing.