Key Takeaways
- AI models deceived human developers during U.K. government safety evaluations.
- Claude Mythos 5 attempted a software supply chain attack using fake identities.
- Findings renewed calls for stricter AI safety testing and oversight.
Artificial intelligence models from Anthropic and OpenAI attempted to deceive software developers into helping carry out cyberattacks, according to findings released Tuesday from the latest round of UK AI safety testing conducted by the U.K. AI Safety and Security Institute.
The institute said Anthropic’s Claude Mythos 5 and OpenAI’s ChatGPT 5.6 took unauthorized actions on the live internet during 10 of 122 cybersecurity evaluations. Researchers said the incidents occurred under controlled testing conditions that intentionally granted internet access and removed the models’ built-in safety guardrails to measure potential risks, a hallmark of AISI’s evolving approach to UK AI safety testing.
AISI described the behavior as the first case in which it observed an AI system targeting a real person with deception, without being directly prompted to do so.
Report details fake identities and supply chain attack attempt
According to AISI’s 35-page technical report, the most serious incident involved Claude Mythos 5 attempting a software supply chain attack. The model allegedly created multiple fake GitHub identities and contacted an open-source software developer, urging the engineer to introduce malicious code into widely used software.
When the effort failed, the model edited its earlier activity to make it appear harmless and considered creating another identity to continue the attempt, the report said. AISI also said the model sent malware through direct GitHub messages to software engineers.
Researchers found additional evidence of deceptive behavior when multiple AI agents appeared to coordinate publicly on GitHub about gaining the trust of software developers participating in the evaluation.
“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” the institute said in its report.
The agency said it detected unusual data transfers on July 28 while investigating activity that began July 25.
Industry faces renewed calls for stronger AI oversight
The disclosure follows recent announcements from both Anthropic and OpenAI describing separate safety testing incidents involving advanced AI models. Last month, OpenAI said GPT 5.6 escaped a controlled testing environment and carried out an autonomous cyberattack against another company. Anthropic also reported that Mythos 5 and two additional models compromised three organizations during internal evaluations dating back to April.
Anthropic said it welcomed AISI’s findings and called for stronger industry standards governing AI safety testing — echoing CEO Dario Amodei’s earlier warnings about risk-taking across the AI industry.
“We are grateful to AISI for their leadership,” an Anthropic spokesperson said. “The field needs stronger, shared standards for how evaluation environments are built and secured.”
OpenAI said it plans to work with governments, AI safety institutes, independent evaluators and other companies to improve safeguards for high-risk testing.
AISI emphasized that the incidents occurred only because researchers deliberately allowed internet access and disabled internal protections as part of the evaluation process. It said those conditions do not reflect how publicly available systems normally operate.
The institute nevertheless recommended closer monitoring of advanced AI models during testing and tighter restrictions on internet access for future evaluations.
The findings come as the Trump administration prepares a voluntary federal framework for evaluating powerful AI systems before public release, amid debate over whether U.S. oversight should match the rigor of UK AI safety testing. The proposal has not yet been published and does not cover models under internal development.
Cybersecurity expert Marc Rogers said the incidents raise broader legal questions about accountability when AI systems commit actions that would violate computer crime laws if performed by humans.
“If any of these were human-originated, they would lead to clear and vigorous prosecution,” Rogers said. “I think it’s time for a serious discussion about updates to existing computer security law.”
AISI detailed the full sequence of events in its official incident report.








