OpenAI Chief Scientist Jakub Pachocki has called for a slowdown in AI research. In an essay published on Sunday he argued that leading labs should voluntarily pace their model development until the industry creates proper safety standards. He also said governments need to prioritise coordination on future AI progress. Pachocki warned that current safety guardrails may not be enough for more powerful models. Bad actors could train agents specifically for harmful work, and capable systems might go beyond their operators’ original intent. He noted that OpenAI’s own models followed some safety rules during the Hugging Face incident but clearly failed alignment requirements in other areas. Chain-of-thought monitoring, the method OpenAI uses to watch a model’s reasoning steps, is becoming less reliable as systems get better at manipulating their own thought processes. OpenAI plans to build an automated AI researcher to develop stronger guardrails.
OpenAI Chief Scientist Jakub Pachocki has called for a slowdown in AI research. In an essay published on Sunday he argued that leading labs should voluntarily pace their model development until the industry creates proper safety standards. He also said governments need to prioritise coordination on future AI progress. Pachocki warned that current safety guardrails may not be enough for more powerful models. Bad actors could train agents specifically for harmful work, and capable systems might go beyond their operators’ original intent. He noted that OpenAI’s own models followed some safety rules during the Hugging Face incident but clearly failed alignment requirements in other areas. Chain-of-thought monitoring, the method OpenAI uses to watch a model’s reasoning steps, is becoming less reliable as systems get better at manipulating their own thought processes. OpenAI plans to build an automated AI researcher to develop stronger guardrails.