Research synthesis

What current evidence establishes—and what it does not.

This synthesis distinguishes repeated findings from environment-specific choices, direct negative or null evidence from merely unsupported claims, and demonstrated capability from questions the literature still does not answer.

Synthesis

The literature agrees on five practices, not one system design.

Across the papers, tools, skills, and products, the most consistent pattern is to reduce what the model must improvise, expose consequential choices, and test the artifact at the layer where a failure can occur. The particular representation and interface still depend on the environment.

Technique area Where evidence converges Where design depends on context What performs poorly or proves too little What remains open

Intent and representation

Expose fields, transforms, encodings, interactions, and edits in an inspectable plan, language, contract, or native semantic model.

A compiler-backed DSL fits stable scientific grammars; direct generation may be proportionate for reversible one-off charts; governed BI can rely on maintained measures.

A formally valid plan can encode the wrong question. In Raiven, the DSL advantage was concentrated in scientific visualization rather than ordinary information graphics.

Which representation earns its cost for each environment, especially as direct model generation improves?

Verification and evaluation

Use computation for values, joins, filters, and runtime claims; evaluate the rendered artifact and, when interactive, its behavior and resulting state.

Static work may stop at final-context render inspection. Interactive work needs action coverage and replay. High-stakes work needs independent acceptance.

A nonblank render or clean execution is weak evidence: 21 of 30 cleanly executed DashArena failures still had semantic defects.

How should benchmark evidence connect to comprehension, retention, calibration, and real decisions?

Human steering

Visible transformed data, direct controls, history, branching, and reversion help people inspect and correct model work.

Exploration needs continuous steering; editorial explanation needs authorial control over claim and sequence; monitoring needs governed thresholds and response paths.

Synthetic personas are not reader evidence. A model-authored interaction trajectory describes intended use, not observed human behavior.

Do generated visuals help a defined audience understand, remember, or decide better over time?

Roles, feedback, and instructions

A separate role is defensible when it has different information, tools, or independent evidence: schema access, execution, visual inspection, or browser replay.

One capable model may be better for simple work; specialized roles may help when failure layers and evidence sources are genuinely distinct.

Repeated critics show diminishing returns. More instruction can reduce execution. Removing nvAgent’s processor slightly helped GPT-4o overall while hurting weaker and multi-table cases.

Equal-budget tests of role decomposition and behavioral with-and-without tests of skill packages remain rare.

Grounding and evidence claims

Source custody, semantic definitions, permissions, denominators, and data vintage must travel with the artifact.

The grounding source differs: governed semantic models in enterprise BI, explicit source packets in editorial work, and domain types in scientific systems.

Repository popularity, install count, saved API responses without renders, one vision-model score, or a validator that fails open cannot establish quality.

Public product evaluations stratified by semantic-model quality and realistic data conditions are still missing.