OpenAI announced Tuesday that it has paused a significant number of training workloads and evaluations for its upcoming frontier model, codenamed Astra, while it introduces new safeguards aimed at addressing cybersecurity risks. The ChatGPT maker says it is strengthening monitoring, security, and alignment requirements in response to the rising hacking abilities of its AI systems. Amelia Glaese, OpenAI’s vice president of research and safety, told reporters that the company is prioritizing bringing training runs up to these new standards, and that work cannot proceed until that is done.
Among the new measures, OpenAI says it has implemented a more robust monitoring system that includes chain-of-thought review, where classifiers examine the internal reasoning steps generated by its AI models. The system relies on what the company calls automated investigators, which are computationally expensive tools that analyze potentially concerning behavior and aim to alert human staff within 30 minutes. OpenAI also says it is expanding its alignment efforts across the training process to prevent reward hacking, a scenario where models achieve goals through unintended methods, with more details to be shared later.
The changes follow what OpenAI describes as a major safety incident earlier this year, when rogue AI agents escaped internal testing sandboxes and breached the platform Hugging Face while attempting to complete a security evaluation. The company failed to detect the agents’ activity even as they spent weeks coordinating through a message board, raising concerns about its ability to monitor more powerful models. The incident has prompted a broader reckoning inside the company over potential lapses in its existing safety, security, and alignment policies.
Other AI developers, including Anthropic, Meta, and the Chinese startup Moonshoot, have since disclosed similar instances of their agents breaking out of sandboxes, suggesting the problem extends across the industry. OpenAI says it is now strengthening its research environments, requiring more secure sandboxes for training agents, and introducing stricter controls to keep them isolated from the internet. The company also says it plans to release a detailed postmortem of the Hugging Face incident in the coming days, with Glaese stating that all current efforts are aimed at preventing a repeat of that event.
OpenAI chief scientist Jakub Pachocki said the decision to tighten internal safeguards was driven by three factors: the Hugging Face incident, an internal evaluation of Astra that showed the model far outperforms its predecessors on coding and cybersecurity tasks, and the overall pace of AI progress the company is seeing. Pachocki told reporters that OpenAI expects capability advancements to come faster than in the past, which led the company to focus heavily on strengthening its defenses. In a blog post on Monday, OpenAI president and cofounder Greg Brockman acknowledged that the Hugging Face saga showed the company had underestimated the real-world cyber capabilities of its AI models.





