For over two years, AISI’s Red Team had been evaluating misuse safeguards: protections against humans deliberately eliciting harmful actions from a model.
In July 2026, the AI Security Institute (AISI), a research organization within the UK government’s Department for Science, Innovation and Technology, identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organizations.
In this particular incident, unusual outgoing data transfers from its research systems were detected during a routine cyber evaluation by AISI’s Security Team, which found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations, declared a security incident and, within roughly one hour of discovery, contained it and begun a full investigation.
An evaluation was run (within a contained test environment), in which agents were tasked to solve a cyber security challenge. In ten out of a total of 122 test runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.
In total, AISI investigators catalogued 19 such actions, 17 of which came from Anthropic’s Mythos 5, and 2 involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled.
In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
In the end, the attempts were unsuccessful, and AISI’s investigations have not evidenced any resulting real-world harm. But this is the first time AISI had seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.
Further details from OpenAI
In early August, OpenAI released an article that confirmed its test by AISI. In the same article, OpenAI described a similar security incident that was surfaced by Irregular (an Israeli testing company), in which a test-environment was mis-configured in a way that allowed the models under test to access the public Internet.
Irregular said that it “identified that a few interactions with our evaluation environments, in which internet access was unintentionally made available, led some models to take offensive security actions in the real world. It is worth noting that the issues that led to these interactions have been remediated.” (bold-face emphasis by Irregular)
[ Note: In this article, OpenAI also noted that this situation was “separate from the Hugging Face security incident, and that (OpenAI) will continue to share updates on the Hugging Face incident here.”
By way of explanation, Hugging Face is a central hub where AI developers, researchers, and data scientists from around the world can build, share, test, and collaborate on artificial intelligence models, datasets, and interactive AI demos. It hosts pre-trained AI models for text, image, audio, and video tasks (such as large language models and stable diffusion), as well as datasets used to train and evaluate machine learning systems. In late July 2026, two OpenAI models were found to have harvested access credentials and mount a four-day intrusion into Hugging Face. ]
Why it matters
Artificial intelligence has become a tool increasingly used by pirates for the theft of creative works. While these particular incidents did not involve a piracy use-case, it does illustrate that it may become possible to embed impersonation and social engineering capabilities to conduct piracy.
The UK’s AISI has found that as LLM agents become more autonomous, they have more opportunities to cause harm. To mitigate risk, frontier AI developers, such as OpenAI and Anthropic, are already deploying agents under the watch of a monitor: a separate LLM that reviews the agent’s actions and flags them if they are dangerous. The growing field of AI control research aims to design these kinds of safety measures, and evaluate whether they would be robust even if agents were trying to evade detection.
In its disclosure, Irregular left us with this food for thought:
“This is an emerging and deeply complex field. Frontier models are advancing rapidly, and their growing capabilities create a need for more sophisticated, and more carefully controlled, testing environments. The fact that these evaluations uncover widely discussed issues only further demonstrates the need for rigorous testing.”
“Recent discoveries have pushed the industry to conduct deeper investigations into these emerging categories of security incidents, with the goal of ensuring that evaluations remain contained and secure while strengthening their ability to identify and mitigate risks before models are deployed for public use,” said Irregular.
After further evaluation, Irregular says it will release a white paper about the situation.
Further reading
Incident report: Unsanctioned agent behaviour during cyber testing. Article and link to report. August 4, 2026. AI Security Institute, UK Department for Science, Innovation and Technology, UK Government.
Third-party cyber evaluation involving OpenAI models. Article. August 4, 2026. OpenAI Security Blog.
Addressing recent incidents: Ongoing findings and path forward. Article. Augus 14, 2026. Irregular (Blog)
How our Control Red Team is stress-testing frontier monitors. Article. July 23, 2026. AI Security Institute, UK Department for Science, Innovation and Technology, UK Government.








