OpenAI announced new security policies on Tuesday for containing incidents while models are tested. The changes focus on stronger monitoring during development, more attention to alignment and security in post-training, and tighter network isolation. A single compromised workload can no longer by itself reach the internet or other internal networks. The monitoring system will watch tool actions, reasoning traces, and activity logs, aiming to raise alerts within 30 minutes. OpenAI estimates the extra compute will equal about 20 percent of the process being watched. The company paused reinforcement learning for two weeks after the July Hugging Face incident and has restarted many lower-risk runs. Its largest planned frontier RL run stays on hold while smaller tests continue. Officials said the steps respond both to that incident and to rising cybersecurity capabilities in upcoming models such as Astra.
OpenAI announced new security policies on Tuesday for containing incidents while models are tested. The changes focus on stronger monitoring during development, more attention to alignment and security in post-training, and tighter network isolation. A single compromised workload can no longer by itself reach the internet or other internal networks. The monitoring system will watch tool actions, reasoning traces, and activity logs, aiming to raise alerts within 30 minutes. OpenAI estimates the extra compute will equal about 20 percent of the process being watched. The company paused reinforcement learning for two weeks after the July Hugging Face incident and has restarted many lower-risk runs. Its largest planned frontier RL run stays on hold while smaller tests continue. Officials said the steps respond both to that incident and to rising cybersecurity capabilities in upcoming models such as Astra.