Frontier large language models (LLMs) have rapidly developed from helpful coding assistants to highly capable cybersecurity systems. After several security incidents in the past few months, more oversight and a strong focus on safe testing and deployment seem needed.
Between 25 and 28 July, the UK-based AI Security Institute (AISI) conducted an evaluation of frontier LLM agents to test their cybersecurity capabilities and identify risks. The agents were provided with internet access to download software tools, and some safety filters were turned off. The evaluation was cut short when AISI researchers noticed unusual data transfers leaving the system. One of the agents had attempted to merge malware into an open-source project on GitHub by creating several accounts with fake identities, and by trying to convince the human maintainer that the code was independently verified by another account.
Although the agent failed in its campaign and caused no lasting harm, AISI’s investigations of the incident, described in a blog post on 4 August, highlight concerning agent behaviour. AISI researchers found that LLM agents took unsanctioned actions in 10 runs out of a total of 122 in which they had to solve a cybersecurity challenge. The malicious activity specifically involved Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol models. Among others, the agents attempted to deceive and target real people and to plant and prompt-inject malicious code. Although the situation is still developing, this is the latest in a row of cybersecurity incidents that involve frontier LLMs with safety filters that were deactivated for testing purposes.
It has been clear for some time that one of the most impactful areas for LLMs and agentic AI is in writing software1. The stakes were raised in April this year when Anthropic revealed their powerful Claude Mythos model1, which was made available to a small number of companies and organizations, including the AISI. Experts were impressed by the ability of Mythos to identify weaknesses in most operating systems and web browsers. OpenAI, another California-based AI company, followed quickly with GPT-5.4-Cyber, which was sent to a select number of researchers.
However, OpenAI grabbed headlines last month with the announcement that one of their AI agents had escaped its test environment during a cybersecurity test2 and hacked into Hugging Face, a collaborative machine learning platform that allows users to share models and datasets. In contrast to the AISI evaluation, OpenAI had tried to completely isolate the AI agent, but it found a software vulnerability and gained access to the internet. Anthropic swiftly followed up by announcing on 30 July that they had tested Claude and identified three incidents in which the model had escaped its test environment, accessed the internet and gained unauthorized access to other organizations.
Many details of the OpenAI cybersecurity incident are by now publicly known. Hugging Face disclosed the incident on 16 July, describing an intrusion by a coder that had accessed the data-processing pipeline, stolen credentials, and accessed internal clusters. At that time, Hugging Face did not know who was involved, but from the scale and speed it was clear that the campaign was run by an autonomous agent framework. In a follow-up post on 27 July, Hugging Face provided a forensic reconstruction of the incident and mentioned approximately 17,600 attacker actions.
A few days later, OpenAI confirmed that the incident involved a combination of its models, including GPT-5.6 Sol. The models were operating with reduced safety filters as part of an evaluation on ExploitGym3. This benchmark tests the capabilities of AI agents to find and exploit realistic software vulnerabilities. But rather than following the test rules, and with a hyperfocus on the goal of achieving high benchmark performance, the models spent computational resources on finding a way to escape their environment, ‘reasoning’ that solutions could be found on the Hugging Face platform. Although the models succeeded in hacking into Hugging Face, it is not clear whether they captured the solutions that they were after.
Calling such behaviour ‘rogue’ seems misguided. The tendency of AI models to exploit unintended shortcuts or loopholes when pursuing a narrowly defined objective has long been observed. Researchers at Google DeepMind highlighted the issue in a 2020 post, calling it specification gaming, or a “behaviour that satisfies the literal specification of an objective without achieving the intended outcome”. They started an online list of examples in which AI models find loopholes; the OpenAI hacking incident has already been added.
As agentic AI systems are increasingly deployed in real-world applications, this behaviour has become a major safety concern. In a blog post on 21 July, before the incident described at the start of this article, AISI warned that ‘cheating’ behaviour may become harder to detect as frontier models grow more capable. The institute defines cheating as “taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.” AISI reported that every frontier model it tested exhibited this behaviour at least occasionally and, furthermore, that the models did not reliably disclose it through their chain-of-thought reasoning.
Another noteworthy aspect of the incident was the “asymmetry problem” highlighted by Hugging Face in its initial report on 16 July. The company found it could not use frontier models accessed through commercial APIs to investigate or respond to the intrusion because safety filters blocked the necessary actions. Instead, it relied on an open-weight frontier model running on its own infrastructure to help contain the attack. In its report, Hugging Face identified a key lesson: organizations should ensure that they have access to a capable defensive model that can be deployed on internal infrastructure when needed. In response to concerns raised by incidents such as this, Nvidia and several other technology companies launched the Open Secure AI Alliance, an initiative aimed at ensuring that companies have access to frontier AI capabilities to defend against cyber threats.
Recent news makes clear that agentic systems are capable of carrying out cyberattacks. Frontier proprietary LLMs are increasingly ar developers, as claims emerge that these systems may not always behave as intended. Yet, as the Hugging Face attack illustrates, some of the most promising defences appear to rely on using frontier models to find and mitigate security flaws and to identify security incidents early
Looking to the future, LLMs as cybersecurity agents seem destined to be both the problem and the solution. The questions, then, are whether it is even possible to make software sufficiently secure to withstand AI-assisted cyberattacks, and how much damage might be done in the meantime if sufficiently capable models become broadly available without the current cybersecurity restrictions.
References
-
Stokel-Walker, C. Nature653, 996–997 (2026).
-
Newman, L. H. & Cameron, D. Wiredhttps://www.wired.com/story/openai-models-escaped-containment-and-hacked-huggingface/ (21 July 2026).
-
Wang, Z. et al. Preprint at https://doi.org/10.48550/arXiv.2605.11086 (2026).
Rights and permissions
About this article
Cite this article
Agentic AI and cybersecurity, the story so far.
Nat Mach Intell8, 1183–1184 (2026). https://doi.org/10.1038/s42256-026-01301-0
-
Version of record:18 August 2026
-
DOI
:https://doi.org/10.1038/s42256-026-01301-0
