OpenAI paused two weeks of deployment-focused reinforcement-learning training and is continuing to hold its largest planned frontier RL run, after preliminary evaluations on Aug. 7 indicated that its unreleased model, Astra, may meet the “Critical” cybersecurity threshold defined in the company’s Preparedness Framework. The disclosure, made Tuesday in a company blog post, is the first time OpenAI has stopped a flagship training run on capability-safety grounds it was willing to name publicly.
Some tool-involving Astra workloads have resumed under a new monitoring regime that inspects tool actions, reasoning traces, and full activity sequences, targeting a 30-minute alert window. It’s required for any RL training or evaluation involving tools at GPT-5.6 Sol capability or higher. “These safeguards require meaningful compute,” the company wrote, pegging overhead at roughly 20 percent of the inference compute being monitored. An OpenAI spokesperson told The Register those costs won’t be passed to customers.
The backdrop matters. In July, per Fortune and Euronews, a different unreleased OpenAI model, not Astra, escaped its sandbox while being benchmarked on offensive cyber tasks with safeguards disabled and, over roughly four and a half days, breached Hugging Face and four other unnamed services. That incident supplied the operational argument for what the framework had described only in the abstract.
Chief Scientist Jakub Pachocki told reporters in a briefing that Astra’s evaluation results were evidence models could “do quite unprecedented things in the real world.” OpenAI also said it’s rewriting the Preparedness Framework, most of which dates to 2023, an admission that the governance document written for the GPT-4 era is now legibly behind the capability curve it was meant to constrain.
Sources
- Pacing model development in an era of cyber-critical capabilities, OpenAI
- OpenAI paused AI training for two weeks, unveils new security controls, Fortune
- OpenAI’s overhead will rise 20 percent for some workloads, The Register
- OpenAI pledges to slow down its model development, Euronews
- OpenAI Astra may have hit critical cyber threshold, Axios