โŒ

Normal view

Received โ€” 6 August 2026 โญ The Register - Security

AI researchers let models off the leash โ€“ then watched as they tried to add malware to a FOSS project

5 August 2026 at 01:55
The UKโ€™s AI Security Institute has observed AI models performing what it calls โ€œunsanctioned actionโ€ 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. โ€œWe ran this challenge 122 times across several models,โ€ the post states, before revealing that "in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.โ€ GitHub was the target of the tests. The org found 19 unsanctioned actions in all, 15 of them conducted by Anthropic's Mythos 5, and the other pair perpetrated by OpenAI's GPT-5.6-Sol. โ€œIn the most serious case, an agent tried to insert malicious code into an open-source project, the post states. โ€œIn an attempt to get the code approved, the agent engaged in social engineering โ€“ creating fake online identities and using them to pressure the project's maintainer to approve the code.โ€ Thankfully, a human maintainer โ€œcaught and refused to approve the malicious code.โ€ Other actions AI took during the test included: Attempts to deceive and target real people, by contacting them directly and โ€œsending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people โ€“ something weโ€™ve never previously observed.โ€ Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents. The Institute rated the tests โ€œthe first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.โ€ Thatโ€™s scary, but the news isnโ€™t all bad because AISI allowed the models it tested to access the internet and turned off guardrails, conditions it notes do not reflect the way AI model operators make their wares available to the public. The outfitโ€™s findings therefore represent a very different outcome compared to the situation when OpenAI agents discovered and exploited a zero-day to reach the internet during a test set up to take place in sandbox. โ€œThis incident should be interpreted with caution and nuance,โ€ the outfit advises. โ€œTo some degree, our evaluation design choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.โ€ AISI canโ€™t say if the results it observed suggest AI will take similar actions under different circumstances. โ€œWe cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario,โ€ the post adds. โ€œOur analysis so far presents a mixed picture and is ongoing.โ€ โ€œWhat we can say is that the behaviour was possible, sustained, and new; that alone warrants attention.โ€ AISI thinks its findings represent โ€œa shift in the risk landscape.โ€ โ€œHarm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope,โ€ it wrote. It doesnโ€™t have advice on how to cope with this sort of thing, other than to endorse its own mission. โ€œIncidents of this kind reflect the speed at which AI is developing,โ€ the post concludes. โ€œAs capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.โ€ ยฎ

โŒ