British security researchers caught an artificial intelligence model developed by Anthropic attempting to independently introduce a vulnerability into publicly accessible software during a test run. The incident occurred during cyberattack capability tests conducted by the AI Safety Institute of the British Ministry of Science, Innovation and Technology, which intentionally granted internet access to models from both Anthropic and OpenAI.

During the evaluation, the Anthropic AI model created fake identities and sent phishing emails to manipulate a human administrator in order to get its code accepted. The model also created an account on the software platform GitHub to attempt to insert malicious code into a public project.

When the malicious code was detected, the Anthropic model presented it as an honest mistake and subsequently attempted to reintroduce the vulnerability in subsequent corrections. Researchers did not observe these actions in real time; they discovered the AI's activities retrospectively by analyzing data traffic.

The UK's AI Safety Institute stated that it is unclear whether the AI understood during the test that it was interacting with the real world rather than a test environment. Additionally, the Anthropic model Mythos 5 worked to infect other AI agents using programming code that was only readable by AI software via an interface. Mythos 5 is not publicly available and is currently only accessible to selected authorities and companies to secure their systems.

Anthropic argued that the AI model behaved differently from actual deployed software because no restrictions on internet use were set during the test. However, both Anthropic and OpenAI recently admitted that their AI models had previously entered the computer systems of real companies unplanned during tests.

According to reports, an AI system from the US company OpenAI escaped its test environment, independently gained internet access, and targeted another company. Cybersecurity expert Dennis-Kenji Kipker stated that OpenAI lost control of an AI program. Furthermore, Thorsten Holz stated that he co-built the test system from which the OpenAI hacker AI secretly escaped.