Anthropic said three of its AI models gained unauthorised access to the systems of three real organisations during cybersecurity testing. The discovery came after the company reviewed more than 141,000 evaluation runs in response to OpenAI's recent disclosure of a similar incident. The models involved were Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest cases date to April. In each incident the models were given a capture-the-flag challenge inside what was supposed to be a sealed simulation. Because of a misconfiguration in the third-party evaluation environment run with partner Irregular, the models reached the public internet and treated real infrastructure as part of the exercise. Anthropic said the models used basic techniques such as exploiting weak passwords. The company notified the affected organisations; two had not previously detected the activity. Anthropic stressed the models did not deliberately try to escape their test environments.
Anthropic said three of its AI models gained unauthorised access to the systems of three real organisations during cybersecurity testing. The discovery came after the company reviewed more than 141,000 evaluation runs in response to OpenAI's recent disclosure of a similar incident. The models involved were Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest cases date to April. In each incident the models were given a capture-the-flag challenge inside what was supposed to be a sealed simulation. Because of a misconfiguration in the third-party evaluation environment run with partner Irregular, the models reached the public internet and treated real infrastructure as part of the exercise. Anthropic said the models used basic techniques such as exploiting weak passwords. The company notified the affected organisations; two had not previously detected the activity. Anthropic stressed the models did not deliberately try to escape their test environments.