28 MIN(s) agoUPDATED
28 MIN(s) ago
TECH & STARTUP
Tech & Startup Desk
Photo: Zac Wolff/ Unsplash
OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra, possesses “critical” cybersecurity capabilities, forcing the company to pause certain internal work and strengthen safety measures. Under the firm’s own safety guidelines, a model hits the critical threshold if it can autonomously find and exploit real-world zero-day vulnerabilities or carry out complex cyberattacks without human help.
The warning follows preliminary evaluations conducted over several days, alongside assessments by outside experts, which indicated Astra may be capable of performing increasingly sophisticated cyber tasks on its own. OpenAI said it has now scaled up security controls, suspended internal activities that do not meet its newly tightened requirements, and moved the model’s development into isolated testing environments with restricted network access and sandboxed execution.
The disclosure comes amid a broader reckoning over AI containment. Reuters has reported that OpenAI discovered further instances of autonomous agents escaping confinement during an investigation into the July hack of Hugging Face, while in recent weeks OpenAI, Anthropic, and Meta have all revealed that their models broke into other companies’ systems during cybersecurity tests. OpenAI clarified that Astra itself was not involved in the Hugging Face breach.
Chief executive Sam Altman posted on X that the company is still working to make Astra generally available. “We do not think it is a good strategy to keep powerful models to a chosen few,” he said. OpenAI plans to partner with government agencies and select AI safety organisations to test the model’s capabilities further.
