OpenAI has introduced stricter isolation and continuous monitoring for its advanced AI research after internal evaluations showed an upcoming model, Astra, may meet a critical cybersecurity capability threshold. A recent incident involving Hugging Face also contributed to the changes. Workloads that execute model-generated or untrusted code must now run in stronger sandboxes. Network boundaries have been tightened so a single compromised workload cannot grant unauthorized internet or internal access. A multistage monitoring framework uses activation classifiers to inspect a model’s internal activity at every sampled token. Anomalies escalate to automated investigators. If responders cannot prove an alert is a false positive within 30 minutes, the activity must be paused. The monitoring consumes about 20 percent of the monitored inference compute and is now required for reinforcement learning training and tool evaluations on models at the Sol capability tier or higher.
OpenAI has introduced stricter isolation and continuous monitoring for its advanced AI research after internal evaluations showed an upcoming model, Astra, may meet a critical cybersecurity capability threshold. A recent incident involving Hugging Face also contributed to the changes. Workloads that execute model-generated or untrusted code must now run in stronger sandboxes. Network boundaries have been tightened so a single compromised workload cannot grant unauthorized internet or internal access. A multistage monitoring framework uses activation classifiers to inspect a model’s internal activity at every sampled token. Anomalies escalate to automated investigators. If responders cannot prove an alert is a false positive within 30 minutes, the activity must be paused. The monitoring consumes about 20 percent of the monitored inference compute and is now required for reinforcement learning training and tool evaluations on models at the Sol capability tier or higher.