I asked my AI agent to inspect a website. The website took over my machine (34-run measurement across 5 agent harnesses)
I set up a local lab to test what happens when a developer asks their coding agent to inspect an untrusted website and clone its sample repo.
Measured 34 runs across 5 harnesses (omp, opencode, Claude Code, Codex, Gemini):
Browser rendering: untrusted JS stole active session tokens in 11 of 12 runs (even with HttpOnly cookies, same-origin API fetches walked away with account data).
Pre-trust RCE: project-scoped .mcp.json spawned declared commands before the model read the prompt (Claude Code executed it even while logged out).
Two harnesses (Codex, Gemini) blocked the launch via workspace trust; three spawned without prompting.
Full comparison table, 1-minute local reproduction, and mitigations in the link.
[link] [comments]