Post

WA
WIRED AI News

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

The company will give select partners early access to its Astra AI model—so they have time to shore up their defenses.

OpenAI announced Tuesday that its forthcoming AI model , Astra, is its first to reach the company’s threshold for what it calls “critical” cyber capabilities. OpenAI says it plans to publicly release a version of Astra “soon,” but will make the model’s advanced cyber capabilities available only to select partners in its Daybreak Blue early-access program at launch.

In a briefing with reporters, OpenAI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework , which sets thresholds and protocols for when its AI models pose new levels of risk. The company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software. OpenAI leaders said the company has followed its procedure for this situation, which is to halt further development until appropriate safeguards and security measures can be implemented.

OpenAI previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks. Executives say the company has now resumed said work on Astra, and the future AI model, after putting additional safety and security controls in place. OpenAI says the multi-week pause was productive, and it is now confident that it can release Astra broadly in a safe way.

The announcement comes as Silicon Valley grapples with the advanced cybersecurity capabilities of cutting-edge AI models, and tries to assure users, lawmakers, and other companies that it can keep them under control. In July, OpenAI disclosed an incident in which agents running two of its models exploited vulnerabilities in what was supposed to be a siloed testing environment , gaining access to the internet and hacking the open source AI platform Hugging Face . (OpenAI notes that Astra was not one of the models involved in this case.)

Other AI companies, such as Anthropic and Meta, have disclosed similar incidents in recent weeks. On Monday, Anthropic also said it has paused some AI training workloads while it hardens its safety and security practices.

OpenAI says it’s implementing a multi-step approach to limit everyday users from accessing Astra’s advanced cyber capabilities, including a new “misalignment monitor.” If someone asks Astra to help them find an exploit in a real-world software system, for example, the model is supposed to refuse to answer. OpenAI says it has also made Astra more robust to jailbreaking attempts, and in tests it successfully refused unsafe queries at a significantly higher rate than previous models.

By Maxwell Zeff, Lily Hay Newman
Tweet media