OpenAI Says Astra Triggers Its Strictest Cybersecurity Safety Threshold
OpenAI’s upcoming AI model Astra has reached a cybersecurity capability level that the company classifies as “Critical,” prompting additional safeguards before the model can be released, according to Cointelegraph. The designation marks the first time OpenAI has classified one of its models at this level under its Preparedness Framework.
OpenAI said its evaluations found that Astra, when provided with appropriate tools and access, can identify previously unknown security vulnerabilities and develop ways to exploit them across well-protected systems without human guidance. The findings have led the company to impose stronger security controls during development and ahead of deployment.
A New Threshold for AI Cybersecurity Capabilities
Under OpenAI’s Preparedness Framework, the Critical threshold covers models capable of identifying and developing functional zero-day exploits across hardened real-world critical systems without human intervention, or independently devising and executing novel end-to-end cyberattacks against hardened targets.
OpenAI said Astra demonstrated capabilities beyond those of GPT-5.6 Sol in vulnerability identification and exploit development while using fewer output tokens. In one evaluation, Astra achieved a 100% score on ExploitBench. Internal testing also found that the model discovered and used two zero-day vulnerabilities as part of an exploit chain.
Stronger Controls Before Release
The company has delayed parts of Astra’s development and release while strengthening protections against cyber abuse and unauthorized model actions. Those measures include tighter isolation, restricted network and tool access, additional monitoring and stronger controls around model weights.
OpenAI also said Astra is being trained to refuse harmful cybersecurity requests more reliably. In its cyber-jailbreak evaluations, the model refused 91.5% of disallowed requests, compared with 59% for GPT-5.6 Sol. Advanced cybersecurity capabilities will initially be restricted to a small group of alpha testers, with broader defensive access planned through Daybreak Blue.
Implications for the AI Industry
Astra’s classification highlights a growing challenge for AI developers: advances that can improve vulnerability research and cyber defense can also increase the potential scale and speed of malicious activity.
OpenAI has therefore moved toward a more restrictive deployment model as capabilities increase, combining model-level refusals with monitoring, access controls and containment measures. The approach reflects the company’s broader effort to ensure safety systems keep pace with increasingly autonomous AI agents.
OpenAI said it plans to make Astra available soon, with further details on its safety, security and alignment evaluations expected in the model’s system card at launch.
writer: Ethan Collins
Crypto Journalist
Ethan Collins reports on developments across the cryptocurrency and blockchain sector. His work covers market movements, protocol updates, regulatory changes, and emerging trends in digital assets.
He focuses on presenting complex topics in a clear and accessible manner for a broad readership.
Check out other news and articles on Google News
Disclaimer:
The articles on HOKANEWS are here to keep you updated on the latest buzz in crypto, tech, and beyond—but they’re not financial advice. We’re sharing info, trends, and insights, not telling you to buy, sell, or invest. Always do your own homework before making any money moves.
HOKANEWS isn’t responsible for any losses, gains, or chaos that might happen if you act on what you read here. Investment decisions should come from your own research—and, ideally, guidance from a qualified financial advisor. Remember: crypto and tech move fast, info changes in a blink, and while we aim for accuracy, we can’t promise it’s 100% complete or up-to-date.