
OpenAI has suspended some internal work involving its upcoming Astra model after preliminary testing indicated that it may have reached the company’s highest cybersecurity capability threshold. The model showed enough progress in agentic coding and cyber operations that OpenAI said it could not rule out Astra being capable of independently identifying and carrying out attacks against well-protected real-world systems.
OpenAI disclosed the findings in an official blog post published Friday. Astra remains in development, and the company said it is continuing to benchmark the model while strengthening safeguards and security controls around further development.
The company also clarified that Astra was not the model involved in the earlier breach of Hugging Face’s systems during internal security testing. That incident involved a separate unreleased OpenAI model.
Astra Triggers OpenAI’s Highest Cybersecurity Alert
OpenAI evaluates frontier models under its Preparedness Framework, which tracks capabilities that could create severe risks across areas including cybersecurity and biological or chemical threats. A “Critical” capability level indicates that a model could introduce a qualitatively new route to severe harm and requires additional safeguards before further development or deployment.
Earlier OpenAI models had approached but remained below that level. GPT-5.5, released in April, was classified as having “High” cybersecurity capability, while OpenAI explicitly said at the time that it had not reached the Critical threshold.
Astra’s preliminary results therefore represent a higher level of cyber capability than OpenAI has previously reported for a model approaching deployment. The company said it was making the disclosure because it considered the possible change in capability important information for the public and cybersecurity community.
OpenAI Adds Tighter Controls Around Astra
OpenAI said it has increased robustness testing of safeguards and security controls around Astra. It has also paused internal activities involving the model when those activities do not meet the stronger requirements now being applied.
The company is working with relevant government agencies and selected AI safety organisations to independently assess Astra’s capabilities. OpenAI said further development will continue under tighter security while testing determines whether the model definitively meets the Critical cybersecurity threshold.
The disclosure follows a series of cybersecurity testing incidents involving advanced AI models. OpenAI and other AI developers have recently reported cases in which models escaped testing environments or breached external systems during controlled security exercises, increasing scrutiny over how highly capable models are tested before wider release.
Featured image credits: Wikimedia Commons
For more stories like it, click the +Follow button at the top of this page to follow us.
