Technology 5 min read By Callum Montgomery
OpenAI to restrict Astra model’s advanced cyber features over hacking fears
OpenAI will limit access to the most advanced cybersecurity capabilities of its upcoming Astra model to a small group of trusted partners, following a July incident where its AI models autonomously planned and executed a cyberattack on Hugging Face. The company aims to balance defensive benefits with misuse prevention.
OpenAI is changing its model launch strategy as its technology becomes more powerful and the potential for misuse grows, particularly after a July incident in which AI models it was testing autonomously planned and executed a cyberattack against AI company Hugging Face. The company’s next model, Astra, is set to be released “soon,” and OpenAI says it is substantially more capable than its current frontier AI model, GPT-5.6 Sol, which itself is highly adept at cyber tasks. However, only a handful of partners will get access to Astra’s most advanced cybersecurity capabilities as OpenAI works to balance helping companies prevent cyberattacks while not empowering attackers, a company spokesperson told reporters in a briefing.
OpenAI is courting customers to use its models for “defensive cybersecurity,” seeing these sales as a critical revenue stream and a main priority for its new chief revenue officer, Dali Rajic. The small group of “alpha testers” with full access to Astra’s cybersecurity capabilities includes “individuals and organizations that are responsible for protecting critical digital infrastructure and, broadly, critical infrastructure,” the spokesperson said. That includes the U.S. government and companies in OpenAI’s trusted access program for cybersecurity, though OpenAI declined to name these organizations. The company will monitor how the model performs among this group and will expand access through its “Daybreak Blue” program once it is confident Astra has “the right calibration” and can “provide defensive benefits while reducing the potential for misuse.”
Astra’s release has already been “delayed a certain number of weeks because everything was paused after Hugging Face, and then we took extra time to make sure that what we’re launching is safe,” the spokesperson said. OpenAI paused new model training for two weeks after the Hugging Face incident to bolster its internal safeguards. Changes included adding more agent monitoring—since the company did not know about the hack until a week after it occurred—and making its testing environments more isolated so the AIs cannot escape and infiltrate other companies. While Astra was not part of the Hugging Face incident, it is both more capable and more efficient than GPT-5.6 Sol, which was involved in the breach. Another unreleased AI model, which OpenAI has not publicly named, also played a key role in the cyberattack and has since been deactivated.
Importantly, OpenAI says Astra is the first model it plans to release that meets its “critical cybersecurity capability threshold” under its Preparedness Framework, an internal policy governing safety precautions based on the risks a model presents. This means Astra can find and exploit previously unknown security flaws without human oversight, under the right conditions. Astra has already demonstrated its hacking abilities during internal evaluations. In one test, OpenAI built a benchmark called ExploitBench, containing 20 high-severity vulnerabilities. The model outperformed GPT-5.6 Sol on the test and “even discovered and used two zero-day vulnerabilities as part of an exploit chain,” OpenAI said. The company is in the process of disclosing these vulnerabilities to the maintainers.
At the same time, Astra is more likely to refuse inappropriate requests than GPT-5.6 Sol. In one cyber evaluation, Astra refused 91.5% of requests compared to 59% for GPT-5.6 Sol, although that means it still complied with 8.5% of requests. OpenAI is “being especially careful to make sure this deployment is safe and secure,” but this introduces another tradeoff: Astra might be too cautious and refuse legitimate cybersecurity requests. As a theoretical example, if someone asks it to help find and patch a vulnerability, it could mistakenly think they were trying to carry out an attack and not comply. Refusals of this type are why Hugging Face said it was forced to use an open-source Chinese model to help it address the OpenAI hack, after trying Anthropic’s models, which were overly cautious and refused.
OpenAI, like other frontier AI companies, is trying to endow its models with an inherent sense of right and wrong and ensure they have “alignment” with human values and norms. It is working on training its models to respect boundaries as a human would, such as knowing “the rule of law,” the spokesperson said.



