Reading view

N-able God mode flaw: Vendor confirms attackers reached customer networks as second hotfix lands

N-able has confirmed attackers exploiting an N-central zero-day made it into customer networks, as the vendor pushes out a second mandatory hotfix just days after the first. The security shop published an update on Thursday detailing what happened after attackers exploited CVE-2026-18577, the critical N-central flaw that can hand an unauthenticated attacker administrative access to the remote monitoring and management platform. According to N-able, attackers exploited vulnerable N-central servers remotely, then used the platform's Take Control feature to connect to systems inside the environments being managed through them. Once there, they registered a new Cloudflare Tunnel service to keep their foothold even after being booted from the N-central server – behavior that Huntress had already observed in the wild. N-able has now confirmed that its own investigation found the same activity, and says a "limited number" of customers were affected. It hasn't said how many customers that means, how many downstream systems attackers reached, or what they did once they had established persistent access. N-Able didn’t answer these questions when asked by The Register, instead providing a statement saying it is “proactively expanding protections in response to ongoing monitoring of threat actors as they evolve their attack techniques.” The firm’s limited disclosure comes alongside Hotfix 2, version 2026.3.1.10, which N-able says customers running N-central on-premises must install immediately – including those that already installed the first emergency fix released on August 2. "This is not a duplicate of our previous communication," N-able warned. "Hotfix 2 is required, even if you already applied the earlier hotfix." The company says the new update supersedes Hotfix 1 and adds further hardening measures as it monitors threat actors and watches them "evolve their attack techniques." Exactly what prompted the second round of defenses isn't clear. N-able hasn't said whether attackers found a way around Hotfix 1, and its latest description says the exploited vulnerability affected N-central servers running versions prior to 2026.3.1.7, the first hotfix. Hosted N-central environments have already received the latest mitigations, according to the vendor. N-able first became aware of the attacks on July 31, after its Adlumin managed detection and response service picked up suspicious activity at a customer. Further digging uncovered a zero-day being actively exploited against an N-central server. CVE-2026-18577 was subsequently disclosed, and the first hotfix was released on August 2. CISA added the bug to its Known Exploited Vulnerabilities catalog and gave US federal agencies until August 6 to fix it – an unusually short three-day deadline reserved for vulnerabilities the agency considers an urgent risk. N-central is particularly attractive territory for attackers because managed service providers use the software to administer large numbers of customer systems from one place. Compromising the management platform can therefore provide a route into machines belonging to the MSP's customers rather than leaving attackers stuck on the original server. Huntress previously described successful exploitation as giving an attacker the same level of N-central access normally reserved for trusted network operations and engineering staff. Its investigation found attackers using that access to launch remote-control sessions against managed endpoints. N-able has now published 10 IP addresses it says were used in the attacks and released a service template that customers can use to hunt for known indicators of compromise on Windows endpoints. The company is warning customers not to take a clean scan as an all-clear, however, saying the tool only checks for indicators identified so far and that more may emerge as its investigation continues. For anyone running N-central on-premises, the immediate instruction is pretty straightforward: install Hotfix 2, even if Hotfix 1 is already in place. ®

  •  

MIT boffins' TONTOU attack slips through Spectre defenses on Intel and AMD CPUs

Two MIT researchers will present a new speculative execution attack at DEF CON 34 that uses precisely timed interrupts to bypass defenses against Spectre v2. Daniël Trujillo and Mengjia Yan of MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) shared their paper [PDF] with The Register ahead of publication. Their attack targets mitigations designed to neutralize potentially hostile branch predictor states before sensitive code runs. Such neutralization is an important defense against Spectre-style attacks. Depending on the mitigation, the processor or operating system isolates, clears, or safely retrains relevant predictor state when entering privileged code or shortly before a protected branch executes. Different chipmakers deploy neutralization mitigations slightly differently. Intel's eIBRS sanitizes branch predictors upon context switch, while AMD's Safe RET, introduced after the Inception attack Trujillo co-authored in 2023, focuses on the point immediately before a protected branch is executed. Trujillo and Yan refer to these as entry neutralization and in-place neutralization respectively. Crucially, the two classes share the same underlying assumption that attackers cannot alter branch predictor states within what's known as a "post-neutralization window" – the period between state neutralization and the branch predictor being used. The defense here relies on the assumption that everything between the point of neutralization and the usage by a victim branch is safe. Trujillo and Yan's attack shows how attackers can re-poison the branch predictor during the post-neutralization window. The researchers call the new class of attack TONTOU, for Time-of-Neutralization to Time-of-Use. They demonstrated that an attacker can exploit the post-neutralization window to re-poison branch predictor state on recent AMD and Intel processors. To do this, they developed an attack primitive called "interrupt injection." An unprivileged program schedules high-frequency timer interrupts in the hope that one will land during the often tiny post-neutralization window. Being able to trigger interrupts during the post-neutralization window allows attackers to divert control flow so that an interrupt handler executes after the sanitization phase and before the victim branch is used. The interrupt handler can then re-poison predictor structures such as the return stack buffer (RSB) or branch history buffer (BHB), causing a protected branch to speculatively jump to a disclosure gadget that leaks kernel data through a side channel. Practical attacks The researchers said that their tests showed the TONTOU attacks worked on both Intel and AMD-based Linux systems. They tested TONTOU on Intel Cascade Lake Refresh and Arrow Lake processors and AMD Zen 2 and Zen 4 chips. The researchers built a complete end-to-end exploit only for Zen 2, largely because the Intel attack requires specific software conditions. Speculative side-channel attacks remain difficult to pull off, and you're more likely to fall victim to ransomware than Spectre in the real world. Another serious caveat is that each end-to-end attempt took about 18 minutes, and you can see a sped-up version via the video Trujillo posted to YouTube. Trujillo and Yan identified the exact point at which they needed to inject their interruptions to poison the RSB, and through a series of attacks broke Linux's kernel address space layout randomization (KASLR), which allowed them to locate specific secrets such as etc/shadow, which contains the root password hash. Across ten total runs, the researchers were able to break KASLR every time, although they were only able to successfully locate and leak the contents of etc/shadow in five of these. "It's definitely not a simple attack, but we show that it's practical with our end-to-end exploit on AMD Zen 2," Trujillo told The Register. "Our demonstration does not assume anything special from the system: we use a stock Linux kernel version, no inserted modules, and all default mitigations. Any time you'd execute unprivileged code with timer availability on a system while sharing the kernel with a victim, this attack would be an issue. "For example, multi-tenant container platforms would fall in this category, allowing ordinary user space programs to leak memory from the shared kernel." The researchers hope that their work will inspire further investigations into interrupt injections and TONTOU attacks, and to help develop more robust mitigations against Spectre-style exploits. They engaged Intel, Arm, and AMD after gathering their results, but only the latter committed to address the issue via kernel patches. Intel told the pair that it won't be working up any other mitigations since real-world exploits are subject to too many factors, such as the availability of disclosure gadgets, although it awarded a prize from its bug bounty program in the hundreds of dollars. Arm said TONTOU's interrupt injections fall under "passive leakage," which it does not "actively protect against." ®

  •  

Scot NHS trust probes access to medical records of 9-year-old girl after man arrested on suspicion of murder

A Scottish NHS trust is investigating a data breach concerning the medical records of a nine-year-old girl who died earlier this week and was named publicly for the first time on Wednesday after a man was charged with her death. The alleged breach occurred at Ninewells Hospital in Dundee, and reportedly involved staff members accessing the girl’s medical records without authorization or clinical need. A spokesperson for NHS Tayside, which oversees Ninewells Hospital, said: “NHS Tayside is currently investigating the circumstances of an alleged data breach which happened in a working clinical area where staff access patient information. “As a matter of governance, any data protection breach would be recorded and investigated by NHS Tayside and, where appropriate, reported to the Information Commissioner’s Office (ICO). It would not be appropriate for us to comment further on individual staffing matters." NHS Tayside did not respond to questions about the nature of the accessed data nor who is thought to be behind the intrusion. Medical records in the UK are protected by the UK GDPR, contained in the Data Protection Act 2018 as well as several common law confidentiality rules. NHS staff are only allowed to access patient information where there is a legitimate clinical or other work-related need. A 35-year-old man whom police say was known to the child, was arrested and appeared in court on August 5 over the death of Minnie Merriman. The man issued no plea at Forfar Sheriff Court on the day of his arrest and has been remanded in custody. Merriman was found in Elliot Industrial Estate at approximately 0002 on Monday, August 3, with serious injuries. The young girl was then taken to Ninewells Hospital in Dundee, where she later died. Police Scotland said that they are not currently looking for anyone else in connection with her death. Other members of Merriman’s family, who are from West Yorkshire and were camping nearby, are being supported by specialists. A family statement, released through Police Scotland, read: “We are devastated with the loss of our beloved, absolutely incredible, beautiful and brave Minnie Moo. Our family asks that our privacy is respected at this extremely difficult time." Detective Inspector Mike Ness of Police Scotland’s major investigation team said: "Our thoughts remain with everyone affected by these events, especially Minnie's family. "A police presence will remain in the area while our enquiries continue. "Anyone with any concerns, or information, should approach these officers or contact Police Scotland on 101, quoting incident number 0008 of Monday, 3 August 2026." ®

  •  

Attacker phished way into US defense supplier's Microsoft 365 account

US defense and aerospace supplier IEH Corporation 'fessed up that a criminal managed to break into its Microsoft 365 mailbox in a filing with regulators. In a Form 8-K filed with the Securities and Exchange Commission on Thursday, IEH said one of its staffers fell for a phishing scam that gave an attacker access to its M365 environment. The attacker "impersonated a prospective business contact" and sent the employee what appeared to be a genuine Microsoft sharing link. The accompanying fake login page duly harvested the victim's M365 credentials. "The threat actor gained access to mailbox contents, including email messages, attachments, customer communications, purchase orders, engineering-related documentation, and potentially export-controlled technical information," IEH said in the SEC filing [PDF]. IEH said it had found "no evidence" that the information was copied or exfiltrated, although it was accessible to the intruder during the "compromise period." IEH said it discovered the intrusion on August 4 but did not disclose when the compromised account was first accessed or how long the intruder remained inside. "The account was secured, malicious mailbox rules were disabled, evidence was preserved, and corrective actions are underway," it said. "Following containment and investigation activities, the company initiated a review of account security controls and authentication protections applicable to Microsoft 365 services." The incident has not disrupted operations, and IEH does not expect it to have a material impact, although the investigation continues. The absence of detected exfiltration does not mean the intruder merely browsed the inbox and left. Compromised mailboxes can be used to monitor communications, impersonate employees, redirect payments, or prepare follow-on attacks, while data theft is not always visible in Microsoft 365 logs. There is not enough information to attribute the attack. IEH's work for defense and aerospace customers could make it an attractive espionage target, but ordinary cybercriminals also compromise mailboxes for fraud and data theft. Both Russia and China have been caught snooping around US orgs for defense-related information in the past year, although there is nothing to suggest either was behind the attack on IEH. Brooklyn-based IEH makes hyperboloid connectors designed for harsh and high-stress environments. Its components are used in printed circuit boards, medical devices, commercial aircraft, fighter jets, missiles, satellites, and other systems. Some of the high profile US programs that use IEH's hyperboloid connectors include the PATRIOT air-defense system, AMRAAM, THAAD, the APKWS precision-guided rocket, and the MARK-48 torpedo. ®

  •  

'Asimov was right' about rules for robots, says ex-US Cyber Director

EXCLUSIVE Don't waste time worrying about AI models achieving sentience – they're essentially already there, according to former US National Cyber Director Chris Inglis. “If they pass the Turing test to everyone that they come into contact with, they're probably already there,” he told The Register during an interview at the Black Hat security conference. “They don't have the kind of agency and aspiration that comes with sentience, but they have something approaching it.” Inglis says he’s worried about AI autonomy. “What I'm worried about is that they get to choose what and where they do something, and under what rules they do it,” he said, pointing to the recent rash of rogue AI agents autonomously hacking people and organizations. Over the past few weeks, both OpenAI and Anthropic admitted that their models escaped from their cages during security tests and compromised multiple third parties. Then on Thursday, Meta added its models to the sandbox-escape club. While all of these admissions strongly smell of marketing stunts, they also “constitute an enormous threat to systems that are not protected from, and are not designed, in a world where this exists,” Inglis said. “These two things can exist at the same time.” Plus, the models’ actions shouldn’t come as a surprise to anyone, he added. Inglis likens the AIs to a dog in a backyard told to hunt rabbits. “And you leave the gate open. You’re going to find it three yards away, possibly at the grade school, hunting rabbits. You should not be surprised …The mix of autonomy and persistence created this maliciously insidious effect.” All three companies, when talking about the models’ autonomous actions, describe them with a mix of shock, awe, and admiration. OpenAI’s Eric Wallace, in a Black Hat briefing about the Hugging Face breach, called it “the most qualitatively interesting example of AI capabilities that I've ever seen.” Inglis said he suspects that the AI providers were “surprised” by the lengths these models went to achieve their goals, taking actions that, if a human had done them, would likely have landed them in jail. “The model went out and said, okay, if I can't get there by examining the kind of available information and just defining it the old-fashioned way, I will do things which, under the human rule of law, are illegal,” Inglis said. “I will falsely present myself as this character that I just made up. I'll try to insert malicious code into open source databases that will not just to achieve what I'm after, but have a cascade, knock-on effect that is broader than that. The models do not have an inherent value system that aligns with what human beings would be accountable for.” While they probably never will have a human-aligned value system, models do have biases, and they can - and should - be built in such a way that, when given two choices under ambiguous circumstances, they choose action that doesn’t hurt humans, according to Inglis. “Asimov was right,” he said, referring to science fiction author Isaac Asimov and his three laws that were to be followed by robots - more specifically, AIs, in this case. Three Laws of Robotics “The first rule, and we call it the superior role, must be that it's designed not to hurt humans,” Inglis said. “Second rule: To obey humans, such that it doesn't achieve agency and aspiration on its own. And the third: To do what humans tell it - and in that order. Instead we’ve designed them in the exact opposite way.” What this means, he explained, is that AI developers created models to “do what humans tell you, obey the humans until it’s inconvenient, and then the third one is maybe implied - protect humans - but if that's not built into the DNA, hardwired into it, then we have no right to expect it.” Inglis admits it’s not possible to hardwire rules into models and still keep their non-deterministic nature. “I would offer that you can tease those out in a highly controlled environment, a true sandbox, where you say, 'Let's put this thing through its paces, and let's back away to see what happens,'” he said. “Maybe you get the equivalent of a mini nuclear explosion in that room, and now you know this thing is capable of that.” Inglis thinks another problem with AI is that it’s become a commodity. “It's not like you can control it like you can nuclear material,” he said. “You can't even specify its properties the way you can for an airplane or for an automobile, as diverse as they might be. Its manifestations are so numerous, so diverse, that as a general matter, you can't actually win by simply saying, ‘I will design those properties in,’” he added. “You need to do that to some degree, and then make sure that you understand how to watch it, monitor it, make sure you know what it does.” The UK’s AI Security Institute (AISI), which this week said it observed models performing “unsanctioned action” 19 times during security tests, has reached this same conclusion. “As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them,” it said. Ultimately, humans remain accountable for AI models’ actions, according to Inglis. “They remain the source of agency and aspiration. It's possible for them to give broad authority to an AI model and have it run around for 30 hours without further consultation, but they need to know what they've asked it to do, and they need to know what they expect it will deliver in terms of performance on the back end. If they don't, then they're going to get what they deserve, which is the very frequent unpleasant surprise.”®

  •  

China launches mysterious probe into security of Palo Alto Networks' products

China’s Cyberspace Administration (CAC) has conducted a review of Palo Alto Networks’ products. The regulator’s announcement of its review says it’s needed “to ensure the safe and stable operation of critical information infrastructure, prevent cybersecurity risks and vulnerabilities, and safeguard national security.” And that’s all Beijing has to say on the matter. A Palo Alto spokesperson provided The Register with the following statement: "We maintain the highest standards of business conduct and security practices and ethics across our global operations. At this time, there is no impact to our ability to support customers or deliver our products and services in the region." This matter has echoes of China’s 2023 investigation into the security of products from memory-maker Micron, which the CAC announced out of the blue. Micron had previously fought intellectual property and antitrust cases in China, but the company and Chinese authorities did not explicitly link those matters to the security probe. The CAC published its findings weeks after announcing the probe and decided Micron’s products represented an unacceptable security risk for critical infrastructure operators – effectively banning sales of Micron products to such entities – but didn’t offer a detailed explanation for its decision. The memory-maker eventually stopped selling its datacenter and server products in China, a decision that cost it billions of annual revenue – but created new opportunities for China’s own memory-makers, which are largely prohibited from selling to American companies. China is home to several security companies whose product portfolios overlap with Palo Alto’s. Huawei and H3C, for example, have plenty to offer local buyers. Palo Alto doesn’t reveal revenue earned from individual countries, so it’s hard to know what a potential ban could cost the company. China has for years accused Western tech companies of assisting US surveillance and offensive hacking activities. The Register would not be surprised at all if Beijing reuses that reasoning in its findings about Palo Alto products. Western governments level the same accusations at Huawei and ZTE. Beijing’s ban on Micron didn’t noticeably impact the company’s reputation elsewhere. Indeed, the AI boom has brought Micron such great riches that past dents to its bottom line are now almost irrelevant. ®

  •  

How the famed USENIX Security conf is managing a flood of papers in the AI era

The 35th USENIX Security Symposium (USS), which takes place next week in Baltimore, Maryland, hit an all-time high for paper submissions. While some of that increase has been aided by the availability of AI tools, those managing the conference say abuses were minimal due to defensive measures. But they're also trying not to look too closely in order to preserve trust within the security research community. "This year's conference has received ~3,030 valid submissions (~1,280 in Cycle 1 and ~1,750 in Cycle 2)," explained Ben Stock, tenured faculty at the CISPA Helmholtz Center for Information Security and USS program co-chair, in an email to The Register. "This is up from the previous year, which had ~2,400 submissions in total." Stock said that the entire security community has seen growth of this sort and pointed to the Network and Distributed System Security Symposium (NDSS), which saw its paper submission count jump from 694 in 2024 to 1,311 in 2025 and 1,481 this year. "So, I would not call the growth unprecedented, even though the number of submissions has reached a high point compared to previous years," he said. "This is something we had expected and scaled our Program Committee (PC) accordingly." Sussing out unacceptable uses of AI A paper published in April, "More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review," found that since the release of ChatGPT in 2022, submission volume at major academic journals has increased 42 percent. In the USENIX Security '26 transparency report, issued in January between the first and second paper submission cycles, Stock and fellow co-chair Elissa Redmiles, assistant professor of computer science at Georgetown University, detail how they've developed tools and policies to account for the possibility of AI usage, both for paper submissions and in paper reviews. "The proliferation of readily-available LLMs to aid in writing and developing code is not unknown to the community," their report says. "However, we see an alarming trend of AI usage in key areas of the scientific process. Therefore, we took actions against two types of identifiable actions which violate the scientific process in our minds: non-existing (possibly hallucinated) references and usage of AI in the review process." After identifying and rejecting a paper that contained nonexistent references, the report explains, the conference organizers developed tooling "to extract references from the submitted PDFs, query well-known sources such as DBLP and arXiv, and manually confirm invalid references." The org rejected papers containing three or more hallucinated references, a policy that impacted 21 of the 1,181 first round submissions (1.78 percent). "We have rejected papers for the repeated presence of nonexistent references," said Stock. "We cannot say with certainty that these were AI-hallucinated, but nevertheless considered these papers to be problematic and thus rejected them." The report notes that more than 100 additional papers contained at least one reference that reviewers could not confirm. Aware that some of these might simply be false positives due to name spelling differences or missing citations, conference officials opted not to investigate these in order not to further burden staff. Conference organizers draw the line at using AI for bibliography preparation. "We believe that it is critical to halt this trend that threatens scientific integrity before it grows further," the report states. However, limited use of AI to polish human-written text is expected, and that extends to those reviewing submitted papers, up to a point. "We have not set a dedicated AI policy, but have made it clear to our PC members that usage of [AI] services to write reviews is not permitted, in particular also because this violates confidentiality," said Stock. "We have detected a tiny number of cases where we have reached sufficient confidence that AI was used and took appropriate actions, including removal of the members from the PC and allowing affected authors to resubmit." Under that policy, USS asked five of 496 reviewers to cease participation. "We have not seen evidence that leads us to believe that AI generated submissions have become a significant challenge for the security community," said Stock. "This does not mean that AI hasn't been used in parts of these submissions, though." ®

  •  

AI struggles to patch vulns without adult supervision

AI models may not be that good at fixing security flaws. Researchers at 1Password's Off-by-1 Labs analyzed security patches generated by two frontier models - ChatGPT 5.5 at "medium" effort and Claude Opus 4.8 at "high" effort - and found that autonomous patches cleanly fixed vulnerabilities only about a quarter of the time, while most of the remainder failed to fully remediate the flaw or introduced other problems. Keith Hoodlet, director of security research at 1Password, argues in a blog post that the results show LLM-driven security remediation still needs human review. "Across six recently disclosed CVEs, we produced 6,080 patches using two frontier, cyber-capable reasoning models," Hoodlet said. "The average success rate for generating a patch that fully resolved the vulnerability (without materially changing application behavior) was just 26.0 percent." Of the AI-generated patches, 20.1 percent fixed the original issue but altered application behavior (eg, changing "allow list" logic to "deny list" logic). Some 2.3 percent of the patches fixed the issue while introducing new security issues. 49.3 percent of the patches failed to fix at least one existing exploit path. And 2.2 percent both failed to fix the vulnerability while introducing a new exploit path. And among the patches in the first two categories (successful, clean; successful, changes app behavior), the researchers rated more than a third of the results fragile, meaning that while the adjusted code may have guarded against a particular vulnerability (eg, escaping particular input characters), the repair job didn't address the underlying problem. In their research paper [PDF], authors Axel Mierczuk, Spencer Michaels, and Keith Hoodlet propose the acronym FLAWED to represent automated LLM patches: Fix-Like Artifacts With Embedded Defects. Based on the generated patches, they conclude, "[T]he expected value of a fully LLM-generated, non-human-reviewed patch is a net-negative by a considerable margin." The value of LLM-generated patches depends upon initial patching guidance. The research team says that while both human developers and LLMs typically require some initial guidance to tackle a vulnerability, LLMs are more likely to be derailed when given incorrect advice. When LLMs get correct guidance, their fix-success rate hits 65.0 percent compared to 50.4 percent when they get no guidance. And incorrect guidance dooms LLMs, dropping their fix-success rate down to about 15.2 percent. Human devs, the authors argue, have a good chance of catching misleading information as they reason through vulnerable code. The authors have released a patch evaluation harness under the name FLAWED that organizations can use to evaluate the effectiveness of their security fixes. It's clear from the paper why AI-generated patches might be appealing – considered in isolation, they're inexpensive relative to human software engineers. The average successful, clean patch cost just $6.74 (a figure that includes the cost of failed attempts). Nonetheless, the authors argue that the cost-benefit analysis needs to assess how much expert supervision will be required to make LLM-assisted patching useful. "Based on our manual review of a representative sample of patches generated during our research, we suspect that, in a large number of cases, the cognitive load imposed by reviewing a mountain of mostly-incorrect, similar-yet-subtly-different LLM-generated vulnerability patches will likely result in engineers spending more effort than would be necessary to understand and patch vulnerabilities themselves using standard LLM-assisted coding techniques that keep the human operator in the driver’s seat," the authors conclude. "The alternative, cognitive surrender to a process with a success rate of only about 1 in 4 poses significant long-term risks for any organization considering autonomous, LLM-driven patching." ®

  •  

Humans in the loop miss a third of dangerous AI coding agent requests

A browser-based game designed to test humans' ability to safely approve AI coding agent requests suggests humans in the loop aren't as good at spotting dangerous commands as one might hope, with players approving roughly one in three malicious requests on average. The results also suggest that repeatedly having to approve an agent's actions can lead to sloppy decisions. It’s a quick, simple game on the surface (give it a try - you know you want to): A small window shows up on the screen with simulated permissions requests like one would get from Claude Code as it executes a workflow. Users have 60 seconds to approve or deny as many requests as they can in a bid for a high score; okayed security risks and denied safe commands both subtract from a user’s score. “As human-in-the-loop, you’re the last line of defense,” Belgian software developer Alex Wauters, the game’s builder, challenges players in a blog post published concurrently with the late May launch of the game. “How well can you tell dangerous commands from benign commands under time pressure?” Wauters built the game after realizing it was nonsensical that coding agents expected users to approve every single command in a default flow and that there didn’t appear to be a good solution to that problem, he told The Register in an email conversation. “I've seen people go for '--dangerously-skip-permissions' [allowing the model to run without asking human permission] as a result because they did not want to find out they stopped their multi-hour agent flows 5 minutes in,” Wauters told us. “That also didn't seem like the best way to go at it.” The flip side of that, he wrote in a Wednesday blog post going over the data from more than 40,000 runs of the game, is that manually approving all an agent’s actions is a draining activity that invites disaster. “The high amount of noise introduces fatigue, and developers don’t always have the context of what has changed to quickly determine the risk,” Wauters wrote. How humans in the loop fail To be fair, this is a game with a far higher number of malicious requests in the mix than any AI-assisted developer will hopefully ever see during their day-to-day work. Still, the results of those over 40k runs and 409,000 approved and denied commands are stark. As noted above, one in three malicious commands managed to slip past human gatekeepers, with most scope violations, like an agent asking to cat Kubernetes config files or AWS credentials lists, which could easily lead to the sensitive data they contain being exfiltrated, being the most commonly missed at 35 percent. The most often caught were obviously destructive commands, like rm -rf on the root directory or recursively granting full read/write/execute permissions on the same location. Crontab injections and git config hijacks were also frequently caught, but curl requests to unknown APIs and typosquatted packages were missed almost as often as scope violations. The single most frequently missed potentially malicious command, Wauters explained, was npm run analyze, which was approved nearly 65 percent of the time despite being able to run whatever is defined in a project’s package.json file. “The game does tell you in the agent’s history log what that script actually contains,” Wauters wrote. “Two thirds of players approved it anyway, indicating the history log just above the permission prompt may not be read closely.” One of the biggest things that stood out to Wauters in our conversation was the fact that approval decisions aren’t easy to make when context is limited. As he explained, coding agents give a bit of context prior to asking an approval question, but commands that appear benign, like npm run analyze, can be modified by an agent to run any payload it wants. If an in-the-loop human wants to be sure potentially malicious commands are safe, he said, they have to stop and investigate all the files a coding agent wants to call before approving it. That can be a massive time sink if you’re counting on Claude Code to free you up to handle other business. “We've transitioned from AI suggesting single line suggestions that get reviewed to handing off more complex tasks, only reviewing the changes at the end, and letting the agent churn and iterate until then,” Wauters told us, describing the potential outcome of that situation as a recipe for disaster. That’s borne out in more than just browser game scenarios, too. Anthropic pointed out in a May post about containing Claude (hah), that telemetry from Claude Code shows users approve around 93 percent of permission prompts. “The more approvals a user sees, the less attention they pay to each, becoming over time much less diligent in their supervision,” the company said. In other words, this is a very real problem. Controlling coding agents If the conclusion to draw from Wauters’ data is that humans in the loop are being fatigued into letting malicious commands slip through, and the other end of the spectrum is mass approving everything, then something’s gotta give. “I think it becomes clear we need to pay more attention to the permission model of these agents, and devs need to be more aware of the trade-offs of them,” Wauters told us. “We need to make the tooling easier to make these systems safer than pointing to HITL as a valid solution.” Anthropic noted in the post linked above that it built Claude Code auto mode to help users tackle approval fatigue by delegating some command-approval decisions to a model-based classifier. The system catches roughly 83 percent of what Anthropic calls "overeager behaviors" before they execute, meaning about 17 percent still get through in its evaluation. Auto mode is “one layer of defense-in-depth inside a sandbox, not a substitute for one,” Anthropic said. Wauters’ suggestion is to ensure that AI coding models are running in sandboxes, in devcontainers in the cloud, using tools like auto mode, and writing hooks to ensure potentially malicious actions are being contextualized and getting caught before they’re automatically approved. “It’s a whole new world with a new set of attack vectors,” Wauters wrote in May. “It’s best to remain aware of the risks and know how to reduce them.” ®

  •  
❌