System · 2024

MatPlotAgent

A plotting agent evaluated with MatPlotBench that adds execution, debugging, and rendered visual feedback.

Open the primary source ↗

Execution-and-critique agents

Generate Matplotlib figures from natural-language tasks and revise them using execution and visual evidence.

Execution and ablation results over 100 author-built benchmark tasks with limited human inspection.

Execution, debugging, and rendered feedback can help on a bounded plotting benchmark.

An author benchmark and model judge do not establish general chart quality; some models regressed under the loop.

Where this appears in the research.

Audience routes that point here.