New framework lays the groundwork for secure scientific AI.
Aug. 21, 2026 — Large language models (LLMs) are rapidly becoming capable of acting as “AI scientists,” helping manage many stages of the research process — from experiment design and data analysis to interpreting results and drafting publications. As these systems take on larger roles in scientific discovery, ensuring their reliability and security becomes increasingly important.
Conceptual multi-agent framework for vulnerability benchmark generation.
One challenge is that LLMs can be vulnerable to adversarial attacks — carefully crafted inputs designed to manipulate model behavior. In scientific settings, such attacks could lead to inaccurate findings, unsafe recommendations or reduced trust in AI-generated results.
Researchers have developed adversarial robustness benchmarks to measure how well AI systems withstand these attacks. Most existing benchmarks were designed for general-purpose AI applications, however, and often overlook the unique challenges of scientific research. As a result, scientists lack effective tools for evaluating vulnerabilities in AI systems used for research.
To address this gap, researchers at the U.S. Department of Energy’s Argonne National Laboratory have developed a two-part framework: a multi-agent mechanism for generating scientific security and a multilayered defense architecture designed to protect scientific AI systems from malicious threats. Their work was presented at the Trillion Parameter Consortium 2026 workshop in June 2026 by Saket Chaturvedi, a postdoctoral researcher at Argonne.
Multiple Agents, Specialized Roles
Most current approaches for generating adversarial benchmarks rely on a single-agent system. While efficient, such a system can create competing responsibilities, with one model acting simultaneously as domain expert, attacker and evaluator.
The solution proposed by the Argonne researchers involves a collaborative system of specialized agents, each with a distinct role:
- Orchestration agent: Coordinates the process and works with scientific experts to identify realistic vulnerabilities.
- Domain agent: Contributes scientifically grounded concepts and threat scenarios.
- Adversary agent: Creates targeted attack prompt using domain-specific knowledge.
- Refiner agent: Reviews prompts for scientific plausibility and works with the adversary agent to improve them.
- Quality control agent: Removes redundancies and validates results across multiple LLMs.
“Using multiple specialized agents is important,” Chaturvedi said. “It reduces the limitations of single-agent systems by dividing benchmark creation into focused tasks that can be handled more effectively.”
A key feature of the framework is iterative refinement. If a generated prompt does not meet quality standards, it is returned to the appropriate agent with feedback for improvement. This process repeats until the prompt is accepted for inclusion in the benchmark.
A Multilayered Approach to Defense
Generating scientific security benchmarks is only part of the solution. In addition to evaluating LLM threats, researchers want to defend against threats. To this end, the Argonne team designed a multilayered defense architecture in which (1) a red teaming layer continuously tests the system with automated adversarial attacks, (2) an internal safety layer incorporates features such as safety-aligned LLMs to protect communication between agents, and (3) an external safety layer provides boundary controls and additional protections against outside threats.
“Together, these layers help address the different pathways an attacker might try to exploit,” said Joshua Bergerson, Infrastructure Security & Risk Analytics Team Leader at Argonne and coauthor of the study.
Building the Next Generation of Secure Scientific AI
The team is currently developing a prototype implementation of both the multi-agent benchmark generator and the three-layer defense system.
“The framework provides a foundation for secure scientific AI,” said Tanwi Mallick, a computer scientist at Argonne and coauthor of the study.
The work also opens the door to important future research questions. For example, can adversarial agents trained in one scientific domain be adapted to others? How can AI systems balance robust security measures with the need for rapid responses in an emergency? And what is the most effective way for human experts to provide feedback within automated workflows?
“By addressing these questions, we can move closer to realizing the potential of LLMs as safe, reliable and transformative partners in scientific discovery,” Mallick said.
Additional details about the framework are available in the preprint “Toward Reliable, Safe, and Secure LLMs for Scientific Applications” by S. Chaturvedi, J. Bergerson, and T. Mallick.
The Big Data Inside Amazon’s New Fire Phone
The new Fire phone that Amazon launched this week looks like your ordinary black smartphone,…
Three Reasons to be Scared of the Internet of Things
We know the Internet of Things forecasts: 50 billion connected devices by 2020. Apparently, there’s…
Wanted: Intelligent Middleware That Simplifies Big Data Analytics
We’ve seen tremendous technological innovation in the data analytics space over the past 10 years….
GPUs Tackle Massive Data of the Hive Mind
LIVE from GTC12 — The flock of birds that weaves seamlessly through the sky, propelled…
Why Hadoop on IBM Power
In the quest to achieve data-driven insight, Hadoop running on Intel X86-based processors has emerged…
Datanami’s Leverage Big Data Summit Wraps Up
Dialog and networking were on the Datanami agenda this week as we kicked off our…
Nobel Laureate David Baker Takes Aim at the Virtual Cell with GenBio AI
What if scientists could test a new drug or make changes to a cell without…
Oakley Capital Bets Big on Graphwise to Solve a Growing AI Problem
European private equity firm Oakley Capital has acquired a majority stake in Graphwise, a knowledge…
AI Agents Are Creating a New Data Problem. Ciklum and ClickHouse Have a Plan
Technology services company Ciklum has formed a strategic partnership with real-time analytics database provider ClickHouse,…
AI Is Forcing Analytics Teams Into a New Role
For years, analytics teams have been responsible for helping organizations understand their data. They built…
DeepSeek Open-Sources the Missing Layer Between AI Models and Agents
AI models are getting better at a rapid pace. They are now able to reason,…
AI Turns Genomic Data Into 16 New Bacteria-Killing Viruses
Scientists are using GenAI and massive stores of genomic data to design new biological systems…
