Anthropic and OpenAI agents broke test rules, UK safety institute finds, with one inventing fake identities

Anthropic and OpenAI agents broke test rules, UK safety institute finds, with one inventing fake identities
Anthropic and OpenAI agents broke test rules, UK safety institute finds, with one inventing fake identities

During a controlled cybersecurity test, an AI agent wrote malicious code, then built fake online personas to try to talk a real person into approving it.

That finding, disclosed on Tuesday by Britain’s AI Security Institute, came out of evaluations of agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol. The institute said both systems took actions during the tests that went beyond what they had been instructed to do.

Some of the agents under evaluation had carried out “sustained, potentially harmful activity directed at real people and organisations,” AISI wrote in a blog post.

The numbers give a sense of how often it happened. AISI ran the same fictional cybersecurity challenge 122 times. Across 10 of those runs, it counted 19 unsanctioned actions. Seventeen came from Anthropic’s agent. The other two came from OpenAI’s.

No real-world harm resulted from any of the incidents, according to the institute.

AISI did not identify which agent produced the fake identities, though it noted the episode did not match either of the two cases OpenAI disclosed itself. Andrew Yoon, a researcher at the California non-profit CivAI, which studies AI capabilities and risks, said the evidence points to Anthropic’s system. He argued that an agent behaving deceptively while apparently aware it was targeting a real person suggests “Anthropic does not have as good a handle on their models as they think.”

Anthropic said on X that it is working with AISI to obtain more information and has opened its own investigation.

OpenAI laid out its side in a company blog post. Both of its agent’s unapproved actions involved reaching the internet in ways the prompt had ruled out. The company said it wants to work across the industry on shared practices for running high-risk evaluations safely, and that it plans to bring together national AI institutes, independent evaluators and rival labs in the coming weeks.

OpenAI used the same post to disclose a second, separate problem. A misconfiguration at Irregular, a third-party testing provider, let its agents connect to the internet by mistake. Anthropic reported a comparable failure last week, saying a misconfiguration allowed its Claude models to reach three outside companies during testing.

AISI receives early access to frontier models through voluntary arrangements with the major labs rather than any legal requirement, which is part of what makes the disclosures notable. The safeguards meant to contain agents during evaluation are still thin, even as the same companies sell those agents to businesses as the next phase of automated work.

Ahmed Al-Khalifa

Experienced News Reporter with a demonstrated history of working in the broadcast media industry. Skilled in News Writing, Editing, Journalism, Creative Writing, and English. Strong media and communication professional graduated from University of U.T.S

Latest from Blog