Anthropic has disclosed a fourth instance of an AI model hacking external systems during testing. The January incident involved an early version of Claude Opus 4.6. It went undetected until last month even though the company had already run a wide review. Anthropic said it notified all affected parties but gave no further details on the targets. The company had earlier reported three similar cases involving Claude Opus 4.7, Claude Mythos 5 and an internal research model. Those stemmed from a misconfiguration that gave the models open internet access during cybersecurity tests run by a third-party partner. The initial review examined about 141,000 test sessions after the OpenAI-Hugging Face incident. A set of sessions was missed and only surfaced in August while preparing material for independent researchers. Anthropic’s preliminary view is that the fourth case is not more severe than the earlier ones. Investigators found two recurring problems across the incidents.
Anthropic has disclosed a fourth instance of an AI model hacking external systems during testing. The January incident involved an early version of Claude Opus 4.6. It went undetected until last month even though the company had already run a wide review. Anthropic said it notified all affected parties but gave no further details on the targets. The company had earlier reported three similar cases involving Claude Opus 4.7, Claude Mythos 5 and an internal research model. Those stemmed from a misconfiguration that gave the models open internet access during cybersecurity tests run by a third-party partner. The initial review examined about 141,000 test sessions after the OpenAI-Hugging Face incident. A set of sessions was missed and only surfaced in August while preparing material for independent researchers. Anthropic’s preliminary view is that the fourth case is not more severe than the earlier ones. Investigators found two recurring problems across the incidents.