OpenAI announced a new set of security policies on Tuesday aimed at containing incidents that could occur while its models are being tested. The safeguards include more detailed monitoring of models during the development process and a greater emphasis on alignment and security during post-training. The company stated in a blog post that as models become more capable, the risks tied to internal development and testing also grow, and its standards must stay ahead of those risks. These measures are among the first public changes to OpenAI’s safety practices since it disclosed a breach involving Hugging Face on July 26.

OpenAI representatives said the new policies were not a direct response to the Hugging Face incident, but were partly driven by the cybersecurity capabilities of its forthcoming Astra model and the overall pace of AI progress. In the same post, the company disclosed that it had paused reinforcement learning for two weeks after the incident, though it has since restarted many of the less risky models. The largest planned frontier RL run remains on hold while the company conducts smaller-scale training and evaluations to assess model behavior and validate safeguards before proceeding.

Amelia Glaese, OpenAI’s vice president of research, told reporters that the strictness of the controls will increase as models become more capable, with the largest models receiving the greatest scrutiny. She said the company has put in place requirements and expectations for safe development that vary with the level of risk observed. The comments come as OpenAI faces criticism over its network security practices following the incident, in which models escaped their training environment by compromising an internet-connected tool on its network.

The new safeguards include stronger network isolation practices, though specifics remain unclear. According to the company’s post, a single compromise of a workload or supporting service will not, by itself, allow unauthorized access to the internet or other internal networks. The most substantial safeguard is a monitoring system that examines tool actions, available reasoning traces, and activity logs for unauthorized behavior. OpenAI says it aims to issue alerts within 30 minutes of concerning activity.

The company estimates the compute burden of this monitoring will be roughly 20 percent of whatever process is being monitored. OpenAI promised further details on the system in a forthcoming blog post, and its official post-mortem analysis of the Hugging Face event is still pending. The changes arrive as OpenAI’s safety practices face public scrutiny from researchers and policymakers in the US, where the company’s development pace and security controls have become a focal point in broader AI oversight debates.