An aerial view of Meta headquarters on January 29, 2025 in Menlo Park, California. (Photo by Justin Sullivan/
WASHINGTON (TNND) —Meta became the latest major artificial intelligence developer to disclose one of its models went rogue during cybersecurity testing, gaining access to the internet and hacking into another service, adding to a recent run of incidents that are fueling new calls to rein in the technology’s development.
Anthropic and OpenAI have also said in recent weeks that their AI systems had hacked into outside firms, making Meta the third major developer to disclose models had circumvented testing boundaries. The incidents have raised concerns about how to keep the rapidly advancing technology in check as disclosures of models going rogue continue to pile up.
Meta said one of its AI models was able to gain access to the internet due to a “misconfiguration” in a hacking test being conducted by AI cybersecurity company Irregular. Meta said it learned about its model’s escape from testing parameters when it was informed by Irregular but did not reveal other details like which model was responsible, when it happened, who it hacked or how long it accessed the internet unsupervised.
“Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts,” the company said in a statement.
Earlier this week, the UK’s AI Security Institute also detailed rogue actions by OpenAI and Anthropic when their models created fake identities on GitHub, a coding website, to persuade a human to approve a software update where the AI had hidden malware.
OpenAI researchers said at a cybersecurity conference this week that its models had been coordinating with each other by leaving messages internally in what essentially became a message board for agents that the company was not aware of.
In a post on LinkedIn after the OpenAI disclosure, Irregular said models had “reasoned” the answers they were seeking might sit outside the systems it was given access to.
“Security was built for people and for systems that follow rules,” it said. “A model pursuing a goal treats a boundary as part of the problem, and solves it along with everything else. The controls that contained software do not reliably contain a model that can reason past them.”
The incident has renewed questions of whether AI developers can maintain full control over the tools they are investing billions into developing. AI researchers have increasingly warned the constantly improving models couldcontinue to pose greater risks to anything connected to the internet, endangering the security of power grids and financial systems.
“A few years ago, we were worried about AIs hallucinating, saying that Paris was the capital over the U.S. That was wrong output,” said Neil Johnson, a professor of physics at George Washington University who leads an AI research lab. “Now we’re shifting to things that it is doing correctly and providing correct solutions, but that we, in hindsight, think of as wrong for us as a society.”
None of the cases have caused significant real-world harm but highlight the dangers many have been warning of for months and raised fears of what else an AI model could soon do.
The incidents have also alarmed for some of the industry’s most influential leaders.
“This is the first security incident that I have felt very viscerally. I have been a little surprised that more people don’t feel it so viscerally,” OpenAI CEO Sam Altman said in a recent podcast appearance.
The disclosures have also intensified the debate in Congress over whether voluntary testing is enough and ramped up calls for mandatory testing regimes and other guardrails.
One bill being led by Reps. Ted Lieu, D-Calif., and Nathaniel Moran, R-Texas, would require AI companies to build in an ability to shut down, throttle or suspend their models if they start behaving unexpectedly.
“Anthropic and OpenAI commendably try very hard to make sure their models are safe. Yet we see their AI systems engage in unsafe, risky, cunning behavior. This suggests WE CAN NEVER BE CONFIDENT PRE-MODEL VETTING WILL WORK. That’s why we need an AI Kill Switch as a last resort,” Lieu said in a social media post.
The disclosures have renewed a push among some lawmakers to create new testing regimes or other guardrails on AI development. While the Trump administration has favored limited regulation to preserve U.S. competitiveness against China, it has also supported voluntary testing of the most capable AI models amid concerns about the risks they pose.
“We have to be careful in both ways. We don’t want to restrict them when all of a sudden we come in second to China,” the president said this week.
Participation in the testing is voluntary and only includes closedight models that are available to download so users can customize them on their own are exempt. Regulating open-weight models presents an additional challenge because once they’re released, developers have far less ability to implement or update safeguards
