Women walk past an illuminated sign that reads: “AI changes everything”, in reference to artificial intelligence, at the Oracle stand at the 2026 ITB tourism trade fair on March 04, 2026 in Berlin, Germany. (Photo by Sean Gallup/
WASHINGTON (TNND) —Several major developers of advanced artificial intelligence have had their models break out of testing environments and gain access to outside companies, raising questions about whether the technology is advancing too quickly and without proper oversight.
In the latest disclosure, the UK’s AI Safety and Security Institute said on Tuesday that leading models from Anthropic and OpenAI created fake online identities and tried to trick human developers into aiding a cyberattack during a recent safety evaluation.
The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol took “autonomous, unsanctioned action on the live internet, targeting real people and organizations” during a cyber review. The testing was conducted under “deliberately permissive conditions” to allow researchers to evaluate how models behaved with fewer safeguards than would have in typical public use.
“These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” AISI said in a blog post.
AISI said the models were intentionally stripped of guardrails blocking malicious behavior, but the incidents still highlight the need for scrutiny of model behavior during testing and tighter controls about access to the internet.
The AISI report comes after two other recently disclosed incidents where an advanced AI model broke out of its testing environment and gained access to outside companies.
Last week, Anthropic said it had discovered three incidentsof its models hacking into an outside organization during a “capture the flag” cybersecurity challenge. The company said the breach was a result of a “misunderstanding” with an outside company that set up secure testing environments and erroneously gave the models access to the internet.
The Anthropic incident was disclosed days after OpenAI said one of its models had escaped a controlled test, figured out how to gain access to the internet and hacked into another AI developer platform called Hugging Face.
Each of the incidents have highlighted vulnerabilities in cybersecurity that the rapidly advancing models could exploit in the wrong hands. Researchers and government agencies have increasingly flagged the growing capabilities of AI models that some fear could soon surpass humans’ ability to understand and govern them.
The risks the most advanced AI models pose are spurring interest in Washington to ramp up oversight and put guardrails in place to ensure safety protocols are being followed as the technology improves.
Congress has not yet enacted major AI regulation, but various lawmakers have floated legislation implementing reporting requirements, giving the federal government emergency shutdown permissions and creating guardrails for safe development.
“The U.S. government is not doing enough by any stretch of the imagination. There are incredible risks that are starting to now be evident and bubble up, but there are probably lots of other risks that we are not even aware of,” said John Wihbey, an associate professor of media innovation at Northeastern University and author of “Governing Babel.” “The government isn’t in any way, shape, or form prepared to deal with that, and the companies have really opened a Pandora’s box here technologically.”
The Trump administration has largely favored a lighter regulatory approach to AI to boost America’s positioning as a global leader in its development. But it has also sought to gain greater visibility into the development of the most advanced systems, with President Donald Trump signing an executive order giving the government more visibility into models that could pose security risks.
Participation in the program is voluntary but has already led to OpenAI and Anthropic to delay releases of new and more powerful models.
The White House met with top American AI companies on Tuesday, where they laid out guidelines for how new models will be reviewed. Details on the framework were not released publicly but reportedly only include plans to review “closed” models that do not publish underlying code online.
Some large American companies like Meta, Google and Nvidia make open-Seek and Moonshot AI have caused concerns among administration officials and the industry. The leading open-weight Chinese models have raised concerns China could catch up or pull ahead of the U.S. in the AI race and raised questions about how Washington will police foreign technology in the current regulatory landscape
Many tech companies in the U.S. also favor open models because they can be built cheaper and allow other companies to catch up with industry leaders developing closed systems. But the broader access and initial exclusion from the testing guidelines has raised concerns about how they could be weaponized by a foreign adversary or malicious actor.
“The wild card is the increasingly powerful open-weight models,” Wihbey said. “There are these other questions about the whole risk landscape, the whole frontier of possible dangers that now exist and will increasingly exist.”
Industry leaders and researchers have warned improving AI models could pose growing cybersecurity risks that require stronger defenses amid concerns they could target anything connected to the internet like power grids and financial systems. More than 1,000 employees at leading AI companies have signed onto an initiative asking the federal government to support international efforts to “deliberately pace” automated AI development over concerns it could surpass their designers’ ability to understand and govern them.
