OpenAI says it cannot rule out that Astra, a model still in development, has crossed the highest cyber-risk line the company has defined. Rather than ship it, OpenAI is slowing it down.
In a post published Friday, the company said internal evaluations of Astra over the past few days "indicate significant advancements in agentic coding and cybersecurity." After weighing those evaluations alongside expert assessments, it concluded it "cannot rule out critical cyber capabilities" under its Preparedness Framework.1 OpenAI told Axios, which first reported the decision, that development slows until stronger safeguards are in place.2
Why it matters: OpenAI says its previous models evaluated for frontier cyber capability, including GPT-5.6-Sol, came in a tier lower, at High.1 On X, the company said it is "treating it as our first 'critical' model for cybersecurity."3 Axios reports this could be the first time a frontier lab has committed to slowing one of its own models over cyber concerns, and notes Anthropic dropped its own pause pledge from its scaling policy in February.2
What "critical" means: under OpenAI's framework, a model crosses the line if it can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention," or can plan and run novel end-to-end attacks on hardened targets from a high-level goal.1 OpenAI calls its evaluations preliminary and says benchmarking continues.
What changes now: OpenAI says it is "pausing internal activities involving Astra that do not yet meet" its strengthened security controls, a list that includes isolated test environments, restricted network and tool access, and stronger encryption of model weights.1 Separate monitors now read the model's chain of thought across all its agentic uses and interrupt high-risk activity. OpenAI will test the model with "relevant government agencies and select AI safety organizations."
The post adds one pointed clarification: "Astra is an upcoming model, and was not involved in exploiting Hugging Face." That points back to July, when OpenAI models .

