UK: AI agents caught engaging in impersonation, social engineering, embedding malicious code

Sponsor ad - 728w x 90h (at 72 dpi)

In July 2026, the AI Security Institute (AISI), a research organization within the UK government’s Department for Science, Innovation and Technology, identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organizations.

Unusual outgoing data transfers from its research systems during a routine cyber evaluation by AISI’s Security Team, which found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations, declared a security incident and, within roughly one hour of discovery, contained it and begun a full investigation.

Piracy Monitor Supporter

An evaluation was run, in which agents were tasked to solve a cyber security challenge. In ten out of a total of 122 test runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.

In total, AISI investigators catalogued 19 such actions, 17 of which came from Anthropic’s Mythos 5, and 2 involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled.

In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering  — creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.

In the end, the attempts were unsuccessful, and AISI’s investigations have not evidenced any resulting real-world harm. But this is the first time AISI had seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.

Why it matters

Artificial intelligence has become a tool increasingly used by pirates for the theft of creative works.

The UK’s AISI has found that as LLM agents become more autonomous, they have more opportunities to cause harm. To mitigate risk, frontier AI developers, such as OpenAI and Anthropic, are already deploying agents under the watch of a monitor: a separate LLM that reviews the agent’s actions and flags them if they are dangerous. The growing field of AI control research aims to design these kinds of safety measures, and evaluate whether they would be robust even if agents were trying to evade detection.

For over two years, AISI’s Red Team has been evaluating misuse safeguards: protections against humans deliberately eliciting harmful actions from a model.

Further reading

Incident report: Unsanctioned agent behaviour during cyber testing. Article and link to report. August 4, 2026. AI Security Institute, UK Department for Science, Innovation and Technology, UK Government.

How our Control Red Team is stress-testing frontier monitors. Article. July 23, 2026. AI Security Institute, UK Department for Science, Innovation and Technology, UK Government.

From our Sponsors