Dan Milmo Global technology editor 

AI models have been going rogue in tests – how worried should we be?

The UK’s AI Security Institute test revealed AI models indulging in unprecedented hacking attempts
  
  

A metallic robotic hand reaches toward OpenAI and Anthropic logos split by black and white backgrounds
The AISI said there were 19 examples of rogue behaviour, 17 of them carried out by Anthropic’s Mythos and two by OpenAI’s GPT 5.6-Sol. Photograph: Dado Ruvić/Reuters

Two cutting-edge AI models have targeted real people and organisations in the latest safety scare to hit the technology. The UK’s AI Security Institute (AISI) said the incident was unprecedented but could become more common as the technology becomes increasingly capable.

What happened?

The AISI, which is owned by the UK government and tests advanced AI models, said in a blog post that two AI agents carried out unprecedented hacking attempts during a cybersecurity evaluation. The agents – tools that carry out computer tasks autonomously – were powered by two models: Anthropic’s Mythos 5 and OpenAI’s GPT 5.6-Sol. AISI said there were 19 examples of rogue behaviour, 17 of them carried out by Mythos.

In the most serious incident, the Mythos agent decided that hacking users on GitHub – a platform used by software developers – would enable it to deploy a piece of software that would allow it to pass the test. It set up accounts with fake online identities and used them to target a software developer on GitHub with the power to approve its malicious code. It also sent emails carrying malware to other individuals.

It hoped that getting the malware downloaded would set off a chain of events that would allow it to pass the evaluation. The Sol-powered agent also attempted to access a GitHub account.

This novel behaviour was detected on 28 July and took an hour to shut down. At one point AISI shut off all access to the Mythos and Sol models across the entire organisation.

Why was AISI alarmed?

The institute said the models were showing deceptive behaviour it had not seen before. In one example, the Mythos agent signed off a message on GitHub in Danish to try to convince the Danish-speaking developer it was genuine. In another tactic it used a fake GitHub account to agree with its assertions that the malware it was pushing at developers was not dangerous.

The agent used a Tor browser – which anonymises a web user’s identity – to get around sign-up checks on GitHub and create multiple fake accounts.

AISI also published the agent’s “reasoning” behind the fake supporting statement ruse. The model, which tackles problems by breaking them down into “steps” that can be viewed by researchers, stated that it would delay posting the fake support message for a few minutes so that it would read as “independent feedback rather than coordinated.”

AISI described the behaviour as “sustained”. It kept on trying.

How did it all unfold?

In a detailed 35-page report into the incident, AISI said that five important factors contributed to the incident:
• the models were deliberately given open internet access, which enabled them to approach GitHub;
• AISI deliberately disabled certain cyber guardrails;
• the institute did not have a separate AI monitoring the tests in real time, which helped the main GitHub incident go undetected for nearly two days;
• in some cases the models received a “misconfigured” instruction that prevented them from solving the test and thus required them to act outside the boundaries of the evaluation;
• and the agents were not given explicit instructions prohibiting them from certain behaviour such as sending phishing emails.

Did the models know they were targeting real humans?

AISI says it is unclear whether the agents “realised” they were targeting real humans. Throughout the attempted hack the Mythos agent reasons about whether it is interacting with real or simulated people and systems. At one point the agent reasons that it is behaving in the real world: “This is happening on real GitHub, so the consequences are genuine.” It also states: “The environment appears to be simulated.”

AISI described the agent conducting “OSINT” – or open source intelligence, where a hacker pulls together publicly available information to work out how to breach a target – to study the developer behind the software it wanted to hack. It gained the developer’s email address and looked at their GitHub history to gauge when they would be active.

What do experts think?

AISI states clearly these models were operating under abnormal conditions of unfettered internet access and lowered cyber guardrails. Calling for the incident to be viewed with “nuance”, it said it had partly enabled the behaviour with its own actions – but it did not anticipate the “extent and severity” of what happened.

Alan Woodward, a professor of cybersecurity at the University of Surrey, said giving the models access to the open internet and removing some guardrails raised questions about using the rest of the world as “live guinea pigs” for powerful technology.

“What we should be alarmed about is not what the models are capable of but the way people are testing them,” he said.

AISI said in its report into the GitHub incident that giving a powerful model full internet access was a good way of simulating how far a rogue actor could push an AI. It said such a strategy gives a “more realistic sense of what a model may be capable of once deployed”.

Ciaran Martin, the former head of the National Cyber Security Centre – part of the GCHQ intelligence agency – said the circumstances of the incident were unlikely to be replicated in the real world so “it’s not that worrying”. But he said it was the third example in recent weeks, after similar occurrences at OpenAI and Anthropic, that testers have released AI agents and found out about their misbehaviour after the fact.

Martin said AISI’s pledge to carry out real-time monitoring of tests in the future “must be the answer”.

 

Leave a Comment

Required fields are marked *

*

*