Anthropic published research showing automated systems can fix AI safety problems. Claude-powered researchers tackled 10 different alignment failures. In every case the systems found fixes that improved the target benchmarks. Overall model capabilities stayed intact. The methods also worked on held-out tests and on models up to 4.7 times larger than the ones used in the research loop. One run on deception closed about 85 percent of the measured safety gap. The automated approach beat proposals from experienced human researchers on average within six hours. It tried more than 50 solutions and produced a method using just over 2,000 training examples. That approach was roughly 15,000 times more efficient than Anthropic’s normal production alignment process. The results offer early evidence that automated alignment work can scale, though the systems still rely on human-defined benchmarks and do not yet improve themselves.
Anthropic published research showing automated systems can fix AI safety problems. Claude-powered researchers tackled 10 different alignment failures. In every case the systems found fixes that improved the target benchmarks. Overall model capabilities stayed intact. The methods also worked on held-out tests and on models up to 4.7 times larger than the ones used in the research loop. One run on deception closed about 85 percent of the measured safety gap. The automated approach beat proposals from experienced human researchers on average within six hours. It tried more than 50 solutions and produced a method using just over 2,000 training examples. That approach was roughly 15,000 times more efficient than Anthropic’s normal production alignment process. The results offer early evidence that automated alignment work can scale, though the systems still rely on human-defined benchmarks and do not yet improve themselves.