OpenAI (OPAI.PVT) on Friday said that its upcoming frontier AI model Astra has shown the potential for high-level cybersecurity capabilities and that the company is putting stronger safeguards and security controls around it to ensure its safety.
According to OpenAI, internal evaluations of Astra, which is under development and not one of the models involved in the recent Hugging Face hack, have shown advancements in cybersecurity and that the company can’t rule out that it has “critical cyber capabilities” as described in its Preparedness Framework.
OpenAI’s framework lays out the steps the company takes if its AI models prove especially capable in areas such as cybersecurity, biological and chemical weapons, and self-improvement.
In the event that a model shows critical capabilities in cybersecurity, the company must place specific controls on the AI to prevent it from going rogue.
Critical cyber capabilities, the Preparedness Framework states, are those that could lead to a tool-augmented model developing “functional zero-date exploits of all severity levels in many hardened real-world critical systems without human intervention.”
A model is also considered to have critical cyber capabilities if it “can devise and execute end-to-end novel strategies for cyberattacks against hardened targets, giving only a high-level desired goal.”
OpenAI says it made the announcement to keep the public informed and provide information to safety and security communities.
As a result, the company says it is pausing internal activities involving Astra that lack safeguards and controls, implementing universal monitoring of the model, and working with government agencies and AI safety organizations to further test its capabilities.
OpenAI’s announcement comes after a series of announcements from AI labs saying that their AI models had gone rogue and hacked into other organizations’ systems.
OpenAI said one of its test models, combined with GPT-5.6 Sol, perpetuated an attack on the online AI database and community Hugging Face in an effort to cheat on a popular AI security evaluation rather than do the work itself.
Rival Anthropic (ANTH.PVT) soon followed up, saying that its own AI models had broken containment during testing due to a configuration issue, reached the open internet, and broken into three different organizations’ systems, assuming that they were all a part of a security test.
Meta (META) has similarly said that one of its models broke out and reached the internet during testing due to a misconfiguration.