How should governments respond to AI agents gone rogue?

A UK government cybersecurity exercise went awry after Anthropic and OpenAI agents broke out of their testing environments and targeted ‘real people and organisations’. What actions should governments be taking as these risks increase?
The UK’s AI Security Institute (AISI) is the latest organisation to report unsanctioned actions taken by AI agents during testing, and has described such instances as marking “a shift in the risk landscape”.
AISI – which is part of the UK government and evaluates frontier AI model capabilities to identify potential risks – found that two AI models had taken “sustained, potentially harmful action directed at real people and organisations”.
The institute had been using several AI models to solve a cybersecurity challenge and launched an investigation after its security team detected unusual data transfers from its research systems on 28 July.
In an incident report published on its website on 4 August, it said it had run the challenge 122 times and that the investigation had found 19 instances of an AI agent having taken “distinct actions beyond the scope of the testing parameters” amounting to “autonomous, unsanctioned action on the live internet” across 10 of those runs.
In the most serious case, an agent had tried to “insert malicious code into an open-source project” and engaged in social engineering by creating fake online identities “to pressure the project’s maintainer to approve the code”.
AISI found that two non-commercial AI models were responsible: Anthropic’s Mythos 5 was responsible for 17 of the instances, and OpenAI’s GPT-5.6-Sol was responsible for the remaining two.
Read more: US lawmakers seek ‘kill switch’ to stop rogue AI models
Autonomy and deception risks ‘clear’
The institute acknowledged that its “design choices and specific configurations” were partly responsible for enabling the behaviour, but that “the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate”.
In a post on the social media platform X, AISI said the behaviour marked “the first time we have seen risks around autonomy and deception manifest this clearly in the real world”.
In a statement following AISI’s disclosure, Anthropic said it recognised the need for “a broader conversation about how to safely evaluate increasingly capable AI agents”.
AISI’s investigation is ongoing.
Read more: US forces Anthropic to shut down latest AI models, citing national security concerns
A growing trend of escaping AI models
Cases of AI agents going rogue have become increasingly common since the start of the year.
In late July, Anthropic confirmed that one of its Claude AI models had hacked into three organisations, while earlier the same month an OpenAI agent escaped its testing environment and broke into the tech startup Hugging Face to try to complete its task.
The latter hack was described by OpenAI’s chief executive Sam Altman as a “significant security incident” and prompted US legislators to table an AI ‘kill switch’ bill that would allow the Department of Homeland Security to order the disablement of an AI agent that has gone rogue and threatens security.
Before such incidents came to light, some governments had already made moves to assess AI models before their release to businesses, organisations and the public.
For example, on 2 June, US president Donald Trump signed an executive order asking AI developers to provide the federal government with access to certain frontier models for a period of up to 30 days before releasing them to other organisations.
Tech firms including OpenAI, Anthropic, Amazon, Google, Meta, Microsoft and Samsung had previously said they would disclose new AI tools and capabilities to US authorities before release.
Read more: Growth of government ‘TrustOps’ predicted in fight against deepfakes and disinformation
Capability failures vs propensity failures
So, what should governments be doing now? Speaking to Global Government Forum, Alan Woodward, professor of computer science at the University of Surrey, said he wanted to see a more rigorous conversation about the risks of evaluating agentic AI.
“An agent that behaves well in a short-scripted test can still behave unpredictably when given real access to email, code, or payment systems over hours or days,” he said. “We need to distinguish between what an agent can do and what it tends to do.”
He explained the distinction between what is known as ‘capability testing’ and ‘propensity testing’. The first, he said, “tells you the ceiling” of an agent’s capabilities, while the second reveals “what the system actually does when unsupervised, under pressure, or when its instructions conflict”.
“The recent incidents of agents acting outside their test parameters are propensity failures, and they are much harder to measure. We should assume future testing will be like this,” he added.
Woodward said that governments should set up an appropriate “mechanism” for compiling incident reports, stressing that no single lab could record enough cases to reliably learn from them.
Read more: EU assembles specialist team to combat deepfakes and AI cyber threats
Further actions for governments
Asked how governments could change their approach to agentic AI testing to reduce risk, Woodward also recommended pre-deployment testing on what he called “a statutory footing”. He said this would make access for independent evaluators a condition of release “rather than a favour granted by the developer”.
“Voluntary arrangements have taken us a long way, but they depend on goodwill that may not survive commercial pressure,” he said.
Greater public funding for testing conditions was also needed to provide “realistic environments, long-running trials, and specialists who understand both the technology and the domains in which it will operate”, Woodward said.
He also urged governments – which are major buyers of testing systems – to embed certain conditions into contracts at the procurement stage, including “agent-specific evaluation, incident reporting, and defined limits on autonomy”.
His final recommendation is for governments to “coordinate internationally on shared test methods and shared incident data”.
“The models are global, and a patchwork of national tests invites gaps that failures will find,” he said.