Abstract
Recent agentic systems demonstrate that large language models can generate scientific visualizations from natural language. However, reliability remains a major limitation: systems may execute invalid operations, introduce subtle but consequential errors, or fail to request missing information when inputs are underspecified. These issues are amplified in real-world workflows, which often exceed the complexity of standard benchmarks. Ensuring reliability in autonomous visualization pipelines therefore remains an open challenge. We present TopoPilot, a reliable and extensible agentic framework for automating complex scientific visualization workflows. TopoPilot incorporates systematic guardrails and verification mechanisms to ensure reliable operation. While we focus on topological data analysis and visualization as a primary use case, the framework is designed to generalize across visualization domains. TopoPilot adopts a reliability-centered two-agent architecture. An orchestrator agent translates user prompts into workflows composed of atomic backend actions, while a verifier agent evaluates these workflows prior to execution, enforcing structural validity and semantic consistency. This separation of interpretation and verification reduces code-generation errors and enforces correctness guarantees. A modular architecture further improves robustness by isolating components and enabling seamless integration of new descriptors and domain-specific workflows without modifying the core system. To systematically address reliability, we introduce a taxonomy of failure modes and implement targeted safeguards for each class. In evaluations simulating 1,000 multi-turn conversations across 100 prompts, including adversarial and infeasible requests, TopoPilot achieves a success rate exceeding 99%, compared to under 50% for baselines without comprehensive guardrails and checks.
AI Overview
Our new overview generator adds more detail and page citations
Reliable Conversational Workflow Automation for Topological Data Analysis
Scientific visualization is a critical tool for domain scientists to explore complex datasets, yet many advanced techniques, particularly in Topological Data Analysis (TDA), remain inaccessible due to the steep learning curve of specialized libraries and programming languages. While Large Language Models (LLMs) have shown promise in automating visualization through natural language interfaces, their inherent stochasticity often leads to hallucinations, invalid operations, and a lack of transparency—issues collectively referred to as reliability challenges.
TopoPilot is introduced as a reliable agentic framework designed to automate complex, multi-stage scientific workflows. By moving away from unrestricted code generation and instead utilizing a constrained, tool-based architecture supported by a two-agent verification system, TopoPilot aims to provide a robust solution for researchers who require both ease of use and high operational correctness. The system specifically targets TDA, a field characterized by intricate pipelines involving data preprocessing, feature extraction, and the visualization of derived structures like persistence diagrams and merge trees.
The Challenge of Reliability in AI Agents
Current LLM-based visualization systems, such as VizGenie or ChatVis, typically rely on generating Python code to execute tasks. While flexible, this approach is prone to several failure modes. An LLM might generate code that uses deprecated library functions, or it might produce a visualization that is syntactically correct but semantically meaningless—for example, attempting to apply a persistence simplification to a vector field that doesn’t support such an operation.
The researchers identify six primary failure modes that plague existing agentic systems:
- Clarification failure (F
1
F_1): Failing to ask the user for necessary parameters when a request is underspecified. - Confusion of capabilities (F
2
F_2): Attempting to perform tasks that the underlying software tools do not support. - Invalid parameterization (F
3
F_3): Using out-of-bounds or logically incorrect values for algorithm parameters. - Invalid workflow (F
4
F_4): Composing a sequence of operations that are semantically incompatible. - Goal misalignment (F
5
F_5): Creating a visualization that does not match the user’s intent. - Erroneous execution (F
6
F_6): Runtime errors during the actual computation or rendering phase.
TopoPilot addresses these by implementing a structured “Node Tree” representation of workflows and a dedicated verification agent that inspects every proposed action before it is executed.
Architecture: Two Agents and a Node Tree
The core of TopoPilot is a two-agent model consisting of an Orchestrator Agent and a Verifier Agent. This separation of concerns is a departure from single-agent systems where the same model both plans and executes, often overlooking its own mistakes.
The Orchestrator and Verifier
The Orchestrator interprets the user’s natural language prompt and translates it into a sequence of atomic tool calls. Instead of writing raw code, it uses the Model Context Protocol (MCP) to interact with a predefined set of tools. Each tool represents a specific scientific operation, such as computing critical points or simplifying a persistence diagram.
Once the Orchestrator proposes a workflow, the Verifier Agent takes over. Its sole purpose is to act as a deterministic safeguard. It analyzes the proposed workflow against the user’s chat history and the internal system constraints. It answers a series of boolean questions to validate the structure: Is the input data compatible with the operation? Are the parameters within valid ranges? Does the workflow actually address what the user asked for? If the Verifier finds an issue, it sends feedback to the Orchestrator for correction, preventing the execution of faulty pipelines.
The Node Tree and Type System
Workflows are represented as a rooted tree where each node is a discrete computational step. This structure provides several benefits:
- Immutability: Once a node is created, its parameters cannot be changed, ensuring that the results are reproducible.
- Caching: Intermediate results are saved to disk. If a user wants to adjust the final visualization (e.g., change the color map), the system can reuse the expensive upstream topological computations without recalculating them.
- Type Safety: The system uses an internal type system where each node defines what data it accepts and produces. This ensures that a scalar field operation is never accidentally applied to a tensor field.
Automating Topological Data Analysis
TopoPilot is specifically tailored for TDA, which provides a rigorous mathematical framework for analyzing the “shape” of data. These workflows are notoriously complex, often requiring the user to navigate multiple abstractions. TopoPilot simplifies this by exposing tools for three main data modalities:
- Scalar Fields: Users can compute persistence diagrams, simplify them to remove noise, and visualize structures like Morse-Smale segmentations or contour trees.
- Vector Fields: The system supports Line Integral Convolution (LIC) and the extraction of critical points to understand flow patterns.
- Tensor Fields: It can handle complex symmetric and asymmetric tensors, using specialized visualizations like hyperLIC and degenerate point extraction.
For example, a user might simply say, “Show me the critical points of this molecule dataset, but simplify it first to remove the noise.” TopoPilot will automatically determine that it needs to load the data, compute the persistence diagram, apply a simplification threshold, compute the critical points of the simplified field, and finally render them.
Empirical Evaluation of Reliability
To test the effectiveness of its safeguards, the authors conducted a massive empirical evaluation involving 1
,
000
1,000 multi-turn conversational trials. They used 100
100 unique prompts, ranging from fully specified to intentionally infeasible requests, and ran each 10
10 times to account for the stochastic nature of LLMs.
The results demonstrate the impact of the Verifier and system guardrails:
- TopoPilot (with safeguards): Achieved a success rate of 99.1
%
99.1%, with only a 0.9
%
0.9% failure rate. - Control Trial (no safeguards): When the correctness checks and verification were disabled, the success rate plummeted to 46.8
%
46.8%.
In the control group, the most common failure was a failure to ask for clarification. When given a vague prompt, the unprotected agent would make unfounded assumptions and generate incorrect visualizations. In contrast, TopoPilot consistently engaged in a clarification dialogue with the user until all necessary parameters were defined.
The computational cost was also evaluated. Excluding the time taken for the scientific computations themselves, the agentic overhead (the “thinking” time) was approximately 31
31 seconds per trial, with an average API cost of roughly $
0.06
$0.06. This indicates that the system is not only reliable but also practical for real-world interactive use.
Ensuring Reproducibility and Trust
One of the major criticisms of AI in science is the “black box” nature of its outputs. TopoPilot addresses this by providing a deterministic path from a chat conversation to a standalone, human-readable Python script.
When a visualization is finalized, the user can export the entire workflow. The system generates a script that uses a library-based version of the TopoPilot backend, explicitly listing every parameter choice and operation. This script can be run independently of the LLM, allowing researchers to apply the exact same workflow to new datasets or include the script in a scientific publication to guarantee reproducibility.
Significance and Future Directions
The development of TopoPilot represents a shift in how agentic systems are designed for science. By prioritizing reliability and verification over the raw flexibility of code generation, the system provides a trustworthy tool for researchers.
Feedback from domain experts in chemistry and engineering suggests that TopoPilot’s greatest value lies in its ability to facilitate rapid prototyping and education. It allows scientists to explore their data through advanced topological lenses without needing to become experts in the underlying software libraries first.
While currently focused on TDA, the underlying architecture—the two-agent model, the node tree, and the deterministic verification system—is generalizable. It provides a blueprint for the next generation of AI-assisted scientific tools that are not just capable, but reliable enough for rigorous scientific inquiry.
ParaView-MCP: An autonomous visualization agent with direct tool use
This citation is identified by the authors as the system most closely related to TopoPilot, as both avoid code generation by exposing capabilities through an agentic interface. TopoPilot is framed as an improvement on ParaView-MCP by adding a flexible backend, extensibility, and systematic validation mechanisms.
S. Liu, H. Miao, and P.-T. Bremer. ParaView-MCP: An autonomous visualization agent with direct tool use. In 2025 IEEE Visualization and Visual Analytics, pp. 61–65, 2025. doi: 10.1109/VIS60296.2025.00018
VizGenie: Toward self-refining, domain-aware workflows for next-generation scientific visualization
VizGenie is presented as a state-of-the-art agentic visualization system that relies on code generation. The authors use it as a key point of contrast to motivate TopoPilot’s reliability-focused architecture, which avoids code generation to mitigate runtime failures and subtle semantic errors.
A. Biswas, T. L. Turton, N. R. Ranasinghe, S. Jones, B. Love, W. Jones et al. VizGenie: Toward self-refining, domain-aware workflows for next-generation scientific visualization. IEEE Transactions on Visualization and Computer Graphics, 32(1):1021–1031, 2026. doi: 10.1109/TVCG.2025. 3634655
ChatVis: Large language model agent for generating scientific visualizations
Similar to VizGenie, ChatVis is cited as a prominent example of a system that translates natural language into visualization scripts. It helps establish the context of current agentic visualization systems and highlights the challenges of code-generation approaches, which TopoPilot is designed to overcome.
T. Peterka, T. Mallick, O. Yildiz, D. Lenz, C. Quammen, and B. Geveci. ChatVis: Large language model agent for generating scientific visualizations. In 2025 IEEE 15th Symposium on Large Data Analysis and Visualization, pp. 22–32, 2025. doi: 10.1109/LDAV68558.2025.00007
The Topology ToolKit
This citation is fundamental to the work, as the paper states that TopoPilot’s backend is implemented using The Topology ToolKit (TTK). TTK provides the core topological computation capabilities that TopoPilot’s agentic framework makes accessible to users via natural language.
J. Tierny, G. Favelier, J. A. Levine, C. Gueunet, and M. Michaux. The Topology ToolKit. IEEE Transactions on Visualization and Computer Graphics, 24(1):832–842, 2018. doi: 10.1109/TVCG.2017.2743938
Audio
Audio summaryTwo hosts talking through the paper’s ideas, methods and results.
Similar papers
UFO2: The Desktop AgentOS25 Apr 2025
Exploring Multimodal Prompt for Visualization Authoring with Large Language Models09 Sept 2026
GuardianAgentBench: Where Agents Fail and How to Guard Them23 Jul 2026
Mozi: Governed Autonomy for Drug Discovery LLM Agents04 Mar 2026
SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems11 Jun 2025
Show moreShow less
MASFactory: A Graph-centric Framework for Orchestrating LLM-Based Multi-Agent Systems with Vibe Graphing20 May 2026
MermaidFlow: Redefining Agentic Workflow Generation
MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific
Data Visualization19 Mar 2024
State and Memory is All You Need for Robust and Reliable AI Agents30 Jun 2025
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production18 Sept 2025
