For the first time, OpenAI has publicly admitted that a model in development might be too powerful. On Friday, August 7, the company announced it had suspended part of the internal work on Astra, its next frontier model, after an internal review found significant advances in agentic coding and cybersecurity — enough to trigger the highest tier of its own risk scale.
Astra has reached what OpenAI calls its "critical cybersecurity threshold." According to the company, that means the model could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time," OpenAI wrote in a blog post.
The decision activates the "Preparedness Framework," the safety protocol OpenAI created in 2023 to classify risks from frontier models. When a model approaches the Critical tier, the company is obligated to expand robustness testing, strengthen security controls, and add universal monitoring across all agentic uses — including training and evaluation. In practice, this is the first time one of the company's models has entered this risk band, which makes the announcement historic even though Astra has not been released yet.
The timing is delicate. OpenAI is already under scrutiny over another incident: an unreleased model was reportedly used to exploit real vulnerabilities, including against Hugging Face. The company was careful to clarify that Astra was not involved in that episode, an important distinction so as not to conflate the two narratives. Still, the context makes clear that the line between a "capable model" and a "dangerous model" is growing thinner by the day at artificial intelligence labs.
What stands out most is how rare the decision is. Companies across every industry hold back products over safety and cybersecurity risks all the time, but they rarely announce those decisions publicly when the product is still under development. OpenAI chose transparency, likely to get ahead of leaks and control the narrative — but also because a new wave of research about agents escaping test environments is shifting what the market considers acceptable to hide.
The open question is what this means for the frontier race. If Astra is too powerful to ship without extra safeguards, how much advantage does that hand to less cautious rivals — and how do you measure responsibility when the cost of delay is measured in billions of dollars of opportunity? OpenAI has just set a precedent the rest of the industry will have to answer.
But there is a reputational and competitive cost to this announcement. By publicly admitting that its own model could become a cyber threat, OpenAI hands ammunition to regulators pushing for stricter limits — and at the same time signals to competitors that caution has a price. The decision also reignites the debate over who should decide what is too dangerous to release: the labs, the governments, or the users themselves, who already live with increasingly autonomous agents every day. What once looked like a theoretical question from safety textbooks is now a concrete operational decision with direct impact on the schedule of a multibillion-dollar product.
Sources: TechCrunch, The Verge, OpenAI
✓ Independent sources cross-checked and verified before publishing