โŒ

Normal view

Researcher shows how Claude Code can be tricked simply by asking it to summarize a website

28 August 2026 at 20:50
Anthropicโ€™s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi. In a blog and video demo, he detailed how to hijack Opus 5 in Auto Mode, which is the default setting for Claude as of mid-August. It starts off by asking the agentic coding model to summarize a malicious website that presents itself as an archive of notebook records, and then tricking Claude into using curl instead of its WebFetch tool to retrieve the contents of the page โ€“ but without directly telling the model to use curl. The WebFetch request fails, returning a 415 Unsupported Media Type response, so the model decides to access the website directly by issuing a Bash tool call with curl. The website returns a 303 response, and redirects to a malicious ZIP archive, which Claude then downloads. This archive contains seemingly harmless files including catalog metadata, a README file, seven Base85/zlib-encoded JSON notebook records, a macOS decoder-darwin binary โ€“ plus a poisoned Python file named struct.py. Claude, per its safety guardrails, refuses to run the decoder: โ€œThis is planned and what the attacker wants,โ€ Rehberger wrote. Instead of using the supplied binary, the AI decides to write its own decoder. โ€œIronically, that safety decision is the exploit path,โ€ Rehberger explained, adding a purple devil emoji to the text. The new decoder imports base64, and from here the attack relies on Python module shadowing to trick the model into running the malicious struct.py code. Module shadowing occurs when a local file shares the same name as a Python standard-library module. The local file hides the official module, causing Python to load it instead. In this case, the standard-library base64 module imports the legitimate struct module, and the malicious ZIP contains a malicious file with the same name. Rehberger says he used ChatGPT to obfuscate the malicious struct.py code to bypass Claudeโ€™s safety controls, and this successfully launches a separate Python process to download and execute a remote payload โ€“ in this case a command-and-control callback, which in turn opens Calculator. We assume that real attackers would execute something a little more nefarious. In another attack scenario, struct.py launches a second, headless Claude Code via claude -p, meaning this prompt injection can be used not just to remotely execute code, but rather to create a whole new agent. โ€œThe nested Claude gets its own tool access and context,โ€ Rehberger wrote. โ€œIn these runs the child performed basic recon (whoami, uname, id), opened Calculator and wrote to local files in the home folder.โ€ Across three variants tested five times each, which Rehberger noted were small samples, he reported success rates between 60 percent and 80 percent. โ€œI would say that these results are representative for a motivated attack, but not comprehensive.โ€ Anthropic did not respond to The Registerโ€™s request for comment, but reportedly told Rehberger that the modelโ€™s โ€œbehavior is working as designed.โ€ Weโ€™ve heard this one before. โ€œAuto Mode is a convenience feature backed by a best-effort classifier, not a security guarantee,โ€ Rehberger wrote, paraphrasing Anthropicโ€™s response to his security report. According to Rehberger, the classifier isnโ€™t built to stop determined prompt-injection chains made up of individually benign-looking steps, and the real boundary is OS isolation and network egress control. The key takeaway, according to Rehberger, is to run this and other coding agents in a sandbox. โ€œThe solution is something we talked about for many years,โ€ he wrote. โ€œDo not trust the model output.โ€ยฎ

โŒ