System · 2025

VisEval

The evaluation harness used by nvAgent for database-to-visualization composition.

Open the primary source ↗

Evaluation harness

Evaluate natural-language visualization queries over one or more databases through execution and task-specific correctness checks.

2,524 queries across 146 databases with heterogeneous execution and result checks.

Makes structured and multi-table visualization-query behavior measurable in its declared environment.

Execution and result checks do not establish readable design, organizational semantic fidelity, or human decision benefit.

Where this appears in the research.

Audience routes that point here.