OpenAI has slowed parts of the development of its upcoming artificial intelligence model, Astra, after internal safety evaluations raised concerns about its ability to carry out sophisticated cybersecurity tasks. The company said it could not rule out that Astra had reached a level classified as a “critical” cybersecurity capability, prompting additional safety measures and a pause in some development activities.
The development marks an important moment in the race to build increasingly capable AI systems. While advanced models can help cybersecurity teams identify vulnerabilities and strengthen digital defences, the same capabilities could potentially be misused to discover weaknesses in computer systems and conduct attacks with limited human involvement.
OpenAI’s preparedness framework defines a critical cyber capability as the ability to autonomously identify and exploit severe, real-world software vulnerabilities, including so-called zero-day vulnerabilities, or carry out complex attacks against highly secure targets without human intervention. Astra’s latest evaluations were concerning enough for the company to activate its safety protocols.
OpenAI has not said that Astra has successfully carried out a real-world cyberattack. Instead, the concern is based on what the model demonstrated during internal testing. The company has said it is taking a cautious approach because the potential consequences of releasing a highly capable AI system without sufficient safeguards could be significant.
Astra is still under development and has not been released as a general-purpose public model. The slowdown therefore gives OpenAI additional time to evaluate its capabilities, strengthen security controls and determine what restrictions may be necessary before further development or deployment.
The issue highlights a growing challenge for the AI industry. As artificial intelligence models become better at writing and debugging code, they are also becoming more useful for cybersecurity research. An AI system capable of understanding complex software can potentially help defenders find vulnerabilities faster. But if that capability becomes sufficiently autonomous, it could also lower the technical barrier for cybercriminals.
That dual-use nature makes AI cybersecurity particularly difficult to manage. A tool designed to help a security researcher identify a vulnerability could potentially be adapted to exploit the same weakness. The difference lies not only in the model’s technical ability but also in the safeguards governing what it can access and what actions it is allowed to take.
OpenAI’s latest decision comes as the company and other AI developers face growing pressure to assess powerful models before they are widely deployed. Traditional software security testing generally focuses on known vulnerabilities and defined attack scenarios. Frontier AI systems introduce an additional challenge because their capabilities can change as models become more capable of reasoning, coding and operating tools.
OpenAI said its recent evaluations showed significant progress in agentic coding and cybersecurity. Agentic AI refers to systems that can carry out multi-step tasks with greater independence rather than simply responding to individual user prompts. That autonomy is one of the reasons cybersecurity researchers are paying close attention to newer AI models.
The concern is not limited to offensive cybersecurity. AI can also become a powerful defensive tool. Security teams can use advanced models to analyse large amounts of code, identify weaknesses, investigate suspicious activity and help develop patches. OpenAI has separately announced work aimed at providing more capable cybersecurity tools to trusted defenders, reflecting the potential benefits of advanced AI in protecting digital systems.
The timing of the Astra decision is also significant because cybersecurity incidents involving AI systems and AI-enabled tools have become an increasing concern. Recent incidents have highlighted how powerful models and software agents can create new attack surfaces, particularly when they are given access to external systems, code repositories or other digital resources.
For businesses, the issue extends beyond the development of one AI model. Companies are increasingly integrating AI into software development, customer service, data analysis and cybersecurity operations. As these systems receive broader permissions, controlling what an AI agent can access and execute becomes an important part of corporate security.
Astra’s evaluation also raises questions about how quickly AI safety frameworks need to evolve. Governments and technology companies are developing rules for testing frontier models, but AI capabilities are advancing rapidly. The challenge is to ensure that security assessments keep pace with improvements in autonomous coding, vulnerability discovery and tool use.
OpenAI’s decision to slow Astra rather than simply proceed with development shows how capability testing can directly affect the release process. The company has indicated that stronger safeguards and security controls will be put in place as it continues evaluating the model.
Astra remains under additional scrutiny. OpenAI‘s internal findings have pushed cybersecurity from being just another capability to a central safety consideration for the model’s development.
The decision sends a broader message to the AI sector: as models become capable of performing increasingly sophisticated technical work, proving that they can be controlled safely may become just as important as demonstrating what they can do.