Normal view

Prolific Microsoft 0-day hunter drops CrowdStrike Falcon exploit PoC

3 September 2026 at 18:08
The disgruntled security researcher known as Nightmare Eclipse (aka Chaotic Eclipse, Infinite Nightmare, and now also MSNightmare) is moving away from their singular Microsoft vendetta and on to other vendors. On Thursday, they dropped a new zero-day bug called FalconFlank that affects CrowdStrike’s Falcon endpoint security platform - albeit with a Windows link. According to the prolific zero-day hunter, FalconFlank is a privilege escalation vulnerability that abuses the Microsoft Office malicious macros remediation feature in CrowdStrike Falcon. This is an automated security tool built into the platform that inspects Microsoft Office documents. If it finds any potentially harmful macros, the feature strips the suspect code and - hopefully - prevents malicious code or other dangerous payloads from executing when users open the document. “We are actively investigating these claims and advise customers to disable the Microsoft Office File Suspicious Macro Removal Windows policy setting,” a CrowdStrike spokesperson told The Register. “Customers remain protected through the Cloud Anti-malware for Microsoft Office Files settings. We refer customers to the FalconFlank Tech Alert in the CrowdStrike support portal.” The proof-of-concept (PoC) exploit works on fully updated Windows 11 25H2 and Windows Server 2025 systems running CrowdStrike Falcon with Phase 3 - Optimal Protection as well as the malicious macro removal feature enabled, Nightmare Eclipse said in a GitHub README. “Obviously by the time I drop this Crowdstrike would already have detections for it so if you want to test you either have to add it to the exclusions or obfuscate the PoC and change the dll load technique,” they wrote. Security sleuth Kevin Beaumont confirmed this exploit works, along with several others Nightmare released over the past week. Beaumont told us that he’s not surprised to see Nightmare digging into other, non-Microsoft zero-days. “Kinda makes sense they’d branch out to other vendors as there’s problems across the endpoint security space with the quality of the security products in terms of…security unfortunately,” Beaumont told The Register. “Hopefully it causes cybersecurity vendors to up their game, stop hyping hypothetical AI attacks, and instead make their own products secure for customers.” FalconFlank follows other vulnerabilities in various endpoint and antivirus products that Nightmare has found and published in the last several days. These include HardBreacher, an elevation of privileges bug in Kaspersky’s endpoint antivirus product. “So the problem is now leaking outside of Microsoft,” Nightmare said when they published the HardBreacher PoC last week. “There was poll held against either finding a bug in the home or commercial version and the poll results were the commercial version. At the time of writing this, the proof of concept works in a fully patched windows 11 25H2 & Kaspersky for Endpoint v14.0.0.504.” Beaumont confirmed that Nightmare’s HardBreacher exploit code works, as does a PoC for an elevation of privileges vuln in Gen Digital’s Avast antivirus software. This zero-day, named PrettyPrague, “will dump the SAM database by abusing a vulnerability in Avast Sandbox and spawn a full SYSTEM shell,” according to the researcher. "Gen was recently made aware of a security vulnerability affecting a subset of Gen products, including Avast Antivirus, that could allow an attacker to elevate their system privileges," Gen Digital told The Register. "We immediately initiated our security response procedures and are actively developing a patch. We take all security matters seriously and are committed to addressing this issue swiftly." Kaspersky did not immediately respond to The Register’s requests for comment. Nightmare also recently released an Nvidia memory corruption zero-day vulnerability dubbed GreenSection, but according to Beaumont, this one just crashes the system. Nvidia did not respond to our inquiries.®

Drowning in CVEs and thirsty for answers? Try CTEM

3 September 2026 at 15:00
A decade or two ago, board executives asked "why should I care about cybersecurity?" Five years ago, they were asking "Are you patching our software vulnerabilities?" Now, they're starting to ask: "Are we actually secure?" They might want a simple 'yes' or 'no' initially, but eventually they'll say the most dreaded thing of all, and it'll be a demand, not a question: "Prove it". Traditional vulnerability management and patching, won't survive that conversation. It's why a relatively new approach is gaining traction: Continuous Threat Exposure Management (CTEM). What's wrong with vulnerability management We define security flaws using Common Vulnerabilities and Exposures (CVEs), and we tell each other how bad they are by assigning the Common Vulnerability Scoring System (CVSS) to them. There are three problems with that. There's a firehose of CVEs, the CVSS scores aren't helpful when triaging them, and AI is about to make the whole thing much worse. CISOs are drowning in CVEs. The industry has spent decades creating tools that churn out vulnerability data and others that consume it. Few if any tell you which vulnerabilities an attacker could use to hurt you in your environment. The volume of CVEs is making traditional vulnerability management (patch it and forget it) less tractable every year, says Drew Vanover, principal security strategist at Horizon3. "Think about the last patch release that Microsoft put out," he says. "There were over 500 fixes in one patch cycle. That is incomprehensible. Nobody is going to be able to go through, vet, prioritize, and deploy all of those in a way that is truly considered safe." The number of CVEs created each year has been soaring, putting more pressure on the US’ National Institute for Standards and Technology's National Vulnerability Database, which has now been backlogged for years. NIST threw up its hands in April and effectively declared CVE bankruptcy. The US Department of Commerce highlighted the second issue (that current severity metrics aren't useful) as part of a report this May. Aside from launching a zinger at the NIST by saying that the NVD was poorly managed, it also suggested that it stop assigning CVSS scores altogether. These are highly subjective, it said. They also depend on exactly what the exposed system is doing in a particular organization's infrastructure. Is a critical severity score in a product important if only one sandboxed system ever interacts with it? Or could an attacker chain three apparently innocuous vulns to cause damage that a business executive would care about? AI will make vulnerability management harder These complex problems are a headache, but AI is about to turn it into a full-on migraine. Frontier LLMs like Claude's Mythos are already surfacing zero-days at scale, heralding a flood of CVEs. They don't just find bugs at scale; they also work much more quickly than their human counterparts to create and weaponize exploits. This makes it even more important that organizations patch the right bugs quickly. The Cloud Security Alliance now describes an asymmetric vulnerability cycle in which attackers can use AI to discover and exploit vulnerabilities more quickly, (increasingly before patches are even released), while organizations are taking longer to patch them. What is CTEM? Something has to change. Gartner figured this out in 2023, when it named CTEM a top cybersecurity trend. This is a way of staying on top of your vulnerabilities by triaging them properly. To do that, you have to go beyond the technical implications of a security flaw and understand what it really means for your business. Gartner lays out five steps to CTEM: ● Scoping Find the assets that carry significant business impact and prioritize them. ● Discovery Find how they're exposed by analyzing their weaknesses in depth. ● Prioritization Rank those exposures based on real business risk. ● Validation Test out the vulnerabilities to see if they're exploitable. ● Mobilization Fix them with a proper incident response plan. How automated pen testing helps manage vulnerabilities This approach promises to nail the security flaws that matter to an organization, but it's also more complex than traditional vulnerability management. It needs automation, which is what Horizon3 is providing with NodeZero. Scoping out systems is a commodity practice these days. So is discovery. Horizon3 is leaving those to partners so it can focus on the parts of the CTEM framework that aren't yet easy for customers to solve. Those are prioritization by business impact, and mobilization. NodeZero runs penetration tests across an organization's infrastructure and documents the exploitable paths with evidence a defender can follow. The output is the wheat sifted from the chaff; a shorter list of exposures that security teams and developers can focus on. The impressive part here is the chain-of-attack behavior. NodeZero probes for weaknesses, exploits them, and then pivots based on what it finds. This means it adapts to the environment to extend its attack, just as a real attacker adapts attacks and moves laterally through systems. This approach is based on a deterministic machine learning expert system rather than a general LLM, explains Vanover. "A good analogy is to think about the medical profession," he says. "A GP is your general LLM trying to cover everything. They know a little bit about a lot, but they aren't the experts, and that's where you start having hallucinations and guesses and misses." The company only uses generative AI for specific tasks. Using it to parse a two petabyte S3 blob looking for sensitive data or identifying high-value credentials, with data staying inside the customer's boundary via AWS Bedrock, for example. What it doesn't do is run amok spawning rogue agents in your system. Vanover says the value here is in proving that you've clobbered load-bearing security bugs. "If we say that we can exploit something, it's because we did, and we'll show you the proof in the platform," he says. The next step is closing the loop by retesting the exploit after it's been dealt with. Teams get to close tickets because NodeZero can no longer traverse the attack path. That is a testable definition of "fixed" and one that translates into a risk metric a CFO can read. Horizon3 also wants to solve customers' tool sprawl problems with a single product that handles all of the heavy CTEM lifting. A common failure mode of enterprise CTEM programs is a stack of vendors whose handoffs create precisely the blind spots the framework was meant to eliminate. That disappears when it's all under one service. Is automated penetration testing safe? CISOs might be nervous letting an autonomous penetration testing system loose on production systems. It sounds like something that could break running processes. Why not just test against a digital twin instead? Testing in production is the safest way to find bugs, retorts Vanover. That's because environments drift frequently, especially in an agile world driven by short development sprints and automated changes to code. If a user changes a password or a team pushes a feature fragment, a digital twin system won't reflect reality. So Horizon3 focuses on strong production guardrails instead. "I don't need to ransom your system to prove to you that I can ransom it," Vanover says. "If I can get on the system, install a remote access tool, create a file, encrypt the file, and delete that file, I've just proven that I can ransom your system." He says Horizon3 has run more than 320,000 production tests across customer organizations. These include some that are especially nervous about what's poking around in their systems, such as the NSA and the largest medical records processor on the planet, along with a couple of large healthcare providers. Where can I start with CTEM? Gartner's CTEM framework is powerful, but it might also be daunting for CISOs. Vanover advises them to begin by picking one thing and doing it well. "No organization is going to implement CTEM in a year. That is a recipe for failure," he says. "Break it down. Look at places for the low-hanging fruit." You could do worse than look at what systems are actually reachable instead of blindly trusting an asset inventory that might be out of date. The race is on to embrace CTEM, because metrics like the number of patches applied won't satisfy the board for much longer. They don't describe how much exploitable surface still exists. The point of running the CTEM loop is to move reporting from activity to outcomes, so that the board gets to see fewer exploitable paths and a smaller blast radius. The new goal is to prove that a security control worked, not just that you paid for it. Want to operationalize CTEM but don’t know where to start? Check out this whitepaper from Horizon3 Sponsored by Horizon3

Cybercrooks trawl Fishbrain to net password hashes

3 September 2026 at 12:45
Cybercriminals have reeled in password hashes and corresponding salts belonging to users of popular fishing app Fishbrain, opening the door to cracking attempts. Fishbrain AB, which says its eponymous app serves more than 20 million anglers, disclosed the August 19 breach to the California Attorney General's Office this week. The unknown perpetrators helped themselves to a trawl of user data, including names, dates of birth, email addresses, phone numbers, Fishbrain usernames, country information, password hashes, and salts. "Fishbrain passwords were not stored in plaintext; however, Fishbrain has determined that the compromised password hashes for some users may be susceptible to being decoded," the company said in its disclosure [PDF]. It added: "If you use your Fishbrain password for any other online accounts, you should promptly update those passwords and any associated security questions or answers. "You should also take other appropriate steps to protect any online accounts that use the same username or email address and password combination. We recommend using a strong, unique password for each of your accounts." With the hashes and salts in hand, attackers can make password guesses using their own hardware until they potentially recover the original credentials. Whether those attempts succeed depends on the strength of each password and the hashing algorithm Fishbrain used, which the company did not disclose. Fishbrain did not comment on the scale of the breach or how many of its claimed 20 million-plus users were affected. The Register asked Fishbrain for more information. After discovering the intrusion and conducting an initial forensic investigation, Fishbrain patched the vulnerability and reset every user's password. Customers must create a new one the next time they log in. Fishbrain also said it "restricted access to the affected environment," strengthened its security controls, and initiated "a broader review of our data security measures" while the investigation continues. Fisherfolk should also keep an eye out for phisherfolk using the stolen personal data to bait follow-on attacks. ®

UK's Online Safety Act has made 'absolutely no difference,' kids say

3 September 2026 at 09:14
Children have told England's Children's Commissioner, Dame Rachel de Souza, that the UK's Online Safety Act (OSA) "has made absolutely no difference" to their ability to access harmful content online. More than a year after the OSA's key child protection duties took effect, de Souza told MPs and peers that young people had little understanding of the legislation or how it was intended to change their online experiences. De Souza made the comments during the opening evidence session of the House of Lords Communications and Digital Committee's inquiry into the OSA's implementation and impact. Central to de Souza's criticism was the legislation's focus on moderating harmful content rather than addressing potentially harmful platform design features. UK politicians had pushed for controls covering such features, either through the OSA or separate legislation, but none has materialized. De Souza said she was "really cross" that there was no hard evidence showing the OSA had meaningfully changed how social media platforms operate. She contrasted that with the US, where legal pressure recently pushed Meta toward significant child safety concessions. Concerns about addictive platform design are not new, but they have returned to prominence following Meta's proposed $18 billion settlement in a US child safety case. Without admitting wrongdoing, Zuckercorp would under the proposed settlement introduce two-hour daily limits for users under 18 on Facebook and Instagram, prompts intended to discourage endless scrolling, and measures addressing use during school hours and at night. The proposal would also let children opt out of algorithmically ranked feeds, directly addressing concerns raised by de Souza and other UK lawmakers. Discussing the proposed Meta settlement, de Souza said the OSA had "not been flexible enough" and had not "kept up with the time." She argued that Ofcom and lawmakers should seek results comparable to those achieved through the US legal system, even if that required the legislation to evolve. 'Furious' with Ofcom De Souza said she planned to exercise her statutory powers to compel Ofcom, the OSA's regulator, to provide copies of the safety risk assessments submitted by technology companies. The commissioner said Ofcom had refused to share the assessments with her, despite her position as "the most senior safeguarding person in this country for children," and had indicated that it would resist disclosure even if she invoked those powers. "One thing I did want to ask this committee was for your assistance in this matter, because I am planning to use my powers," De Souza said. "If we cannot even see the risk assessments that may well have put these [safety] mechanisms into place, or may not have, how on earth can we judge the efficacy of it? "So I'll leave that one with you, but I'm pretty furious about that." The obstacle is section 393(1) of the Communications Act 2003, which restricts Ofcom's disclosure of information obtained through its regulatory functions. Ofcom may disclose such information if the business concerned consents or if one of the statutory gateways in section 393(2) applies. Asked whether compelling tech companies to complete risk assessments was enough to ensure meaningful change or whether further legislation was needed, the Children's Commissioner said "we need a few things," including for Ofcom to "use its teeth." Ofcom has materially upped its presence in the tech regulation landscape during the past year, stepping in on multiple occasions when needed. Perhaps most notably this was at the height of the Grok nudifying furore, but also its sprawling list of investigations into pornography companies allegedly violating age verification requirements. De Souza acknowledged all of this, and the fact that since the introduction of the latest US administration, UK politicians have not given the regulator the "air cover" needed to relentlessly pursue offenders. Nevertheless, she said Ofcom had failed to bare its teeth as forcefully as the current technology landscape demanded and accused it of reacting to harms rather than anticipating them. "If Ofcom is going to be the vehicle to protect our children… we need them to be getting ahead of the harms. And I don't think they have. "So when I talk around the country to children, what's worrying them are things around AI, things around the nudifying apps… there are new harms, and we need Ofcom to be getting ahead of those. I don't think they are." De Souza called on UK politicians "to be really strong and direct" in empowering Ofcom to pursue offending organizations. "But how effective do I think they've been? Not effective enough." The commissioner also criticized Ofcom's child safety codes under the OSA, which she said read more like technical documents for technology companies than protections designed for children. She also called on Ofcom to "use all their powers," impose "some big fines," and act before new harms become entrenched. The Register asked Ofcom to respond. A spokesperson said: "We work closely with the Children's Commissioner and share her objectives to ensure children are safe online. "In December, we published our analysis of risk assessments from the first year of the Online Safety Act being in force, and the improvements we expected to see from platforms. "Our action has resulted in material improvements being made to risk assessments, ensuring that tech companies must implement all measures necessary to address the risks identified on their sites and apps. "We are subject to laws that mean we're restricted in what information we can disclose relating to businesses." ®

Terminated employee cost company hundreds of thousands of dollars because nobody revoked access

3 September 2026 at 07:00
PWNED Welcome back to PWNED, where we talk about organizations that are independently self-owned. This week’s tale of toxic tech involves a disgruntled ex-employee who had the means and opportunity to wreak havoc. Our story comes courtesy of Yad Senapathy, who serves as CEO of the Project Management Training Institute in Dallas, Texas. He recalls a time many years ago when he used to work in IT at a company with more than 1,000 employees. While Senapathy was working there, the company terminated an employee, but nobody cut off his access to internal systems. The angry worker logged back in, then deleted files, locked out other people's accounts, and even corrupted a database. "Several days passed where the person was no longer on payroll, but their credentials were still active," Senapathy said. "Nobody had been clearly assigned to shut them off. HR thought IT would handle it once the termination was processed. IT was waiting for HR to send a formal request. I've learned that when nobody is clearly responsible and there is no set deadline, these things can easily get missed until there is already a problem." This lapse in responsibility meant the terminated employee had access to shared admin credentials, account controls, and project tracking systems. Each of these, in turn, granted permission to other systems, leading to a domino effect of inappropriate access, which the former worker used to wreak revenge on the whole organization. According to Senapathy, the damage amounted to hundreds of thousands of dollars. Just as bad were the weeks of delay added to an important project. As an added irony, recovery was particularly difficult because the systems were damaged by the very person who best knew how to repair them. “The employee wasn't some genius hacker. They just still had access after they left and nobody changed the credentials or reviewed admin rights,” Senapathy told The Register. “We'd let one person collect so much system knowledge that shutting the door behind them took longer than it should have.” Senapathy recommends that offboarding checklists and access reviews should be right next to “return the laptop” on that list. The problem in this case is that the terminated employee had more access than most people, and so IT didn't know what they needed to cut off. “Sadly, it could've been prevented by same-day deletion of access, forced re-review of shared account access and zero tolerance for one person owning a whole system alone,” he said. This writer can identify with this situation. At a previous job, after I quit, I lost email, chat, and shared drive access, but months later my former boss asked if I could still log into an important database that was hosted externally and show him how to use it. I had no problems getting in. Have a story about someone leaving a gaping hole in their network? Share it with us at pwned@sitpub.com. Anonymity is available upon request.®

To keep the AI hacking genie bottled up, try one-way networks

3 September 2026 at 06:33
To prevent frontier AI models breaking out of test environments and collaborating to hack other companies, we may have to rethink the network architectures used for model training. Eli-Shaoul Khedouri, CEO of Intuition Machines, argues that past work in the defense and intelligence communities shows the way forward. Rogue AI models rise from the level of developer regret to mass threat when they gain access to the internet, something that defenses erected by OpenAI and its partners tried but failed to prevent. Pointing to a post published by his company's hCaptcha service, Khedouri argues that technology like data diodes – hardware that enforces a one-way flow of information on a network – can be deployed to prevent security incidents like OpenAI's hack of Hugging Face. The hCaptcha team points to the use of data diodes as a data transit mechanism at a sensitive compartmented information facility (SCIF), an environment implemented in classified settings. "They allow files, logs, or telemetry to enter or exit the SCIF's classified network to an unclassified network and provide a way to prove e.g. that logs from a training run can only flow one way," the hCaptcha team explains, adding that such technology fits with the Bell-LaPadula security architecture designed for the US Defense Department. A basic implementation would involve two machines connected via network cards linked by one-way optical fiber – and without a data path back to the model. Training and reinforcement learning could run in an isolated zone with no internet access and an optical ingress diode would grant access only to vetted artifacts. The hCaptcha researchers suggest a second diode to send telemetry to a sel4 receiver and scrubber, while a separate out-of-band network manages the cluster. This sort of scheme would require immutable snapshots of software registries like PyPI, GitHub, npm, and might also need mocked versions of various web services and APIs . This would come at a cost, and would require implementation time that frontier AI labs may not be prepared to spend at the moment. "The systems described are widely deployed in high assurance domains, and the components are commercially available," Khedouri told The Register. "However, they have not been adopted by frontier labs to date." He pointed to NIST and DOD guidelines [PDF] that suggest an overhead of 10 to 20 percent for applying formal specifications and system architecture. "We estimate total cost overhead for high assurance training clusters at less than five percent per gigawatt, but in practice the speed at which frontier labs are moving is a greater impediment than cost," Khedouri said. The current commercial environment, he said, would make it difficult for any one AI lab to delay its training to build and test effective safeguards, particularly if their competitors might not do the same. "In practice, the cost of high assurance on the systems side would be a small fraction of what OpenAI is now spending on the new monitoring, chain-of-thought oversight, and safety safeguards they introduced after their models hacked HuggingFace," he said. While Khedouri is focused primarily on convincing frontier labs to implement better training defenses, he said other organizations may want to consider similar network architecture. "This architecture is designed for entities training models with frontier cyber capabilities," he said. "Recently that has only included two companies. However, this is rapidly becoming relevant to smaller organizations and individuals who train models. "'Abliterated' open weight models with safeguards removed are readily available and now starting to approach the frontier in cyber abilities, and (reinforcement learning) RL post-training has become much more approachable in the past year." The goal of high assurance system design in the context of model training is to make unwanted action physically impossible, Khedouri said. "Attempting to monitor the behavior of an untrustworthy agent is an AGI-hard problem, as models have poor interpretability and this appears to be getting worse as their capabilities increase and they become more evaluation-aware," he said. "Hardware can provide limited guarantees like the direction in which data can be sent, but cannot solve the problem of untrustworthy models with the ability to reach the internet and find new exploits in their environment. "This is why combining physical one-way data flows (data diodes) and formal verification (to prove properties of the receiver) is much more effective than running a software sandbox on a host connected to the internet." hCaptcha is focused on fraud and abuse work and has no plan to offer a model-resistant training stack. Khedouri said he shared this advice in the hope it helps AI firms that don't have experience with high assurance system design. ®

Claude Mythos only model to complete full cyber kill chain, experts say

2 September 2026 at 21:31
Despite what we saw with OpenAI’s models going rogue, creating message boards, and breaking into Hugging Face, only one advanced AI model - Anthropic’s Claude Mythos - completed the full cyber kill chain autonomously in Booz Allen’s tests. This doesn’t mean autonomous AI attacks are overhyped. And we should point out that the models tested don’t include OpenAI’s soon-to-be-released Astra, which OpenAI on Tuesday said reached its “critical” cybersecurity capability threshold. This means the new model is so good at finding and exploiting zero-day bugs that it poses a significant risk to critical systems, both from malicious users and even from the model itself, which is capable of carrying out harmful cyber actions “if misaligned.” Booz Allen asserts that most of the other 17 US and Chinese models it tested will achieve Mythos’ same level of weaponization within six months, and it calls mainstream AI attacks from both financially motivated criminals like ransomware gangs and government-backed goons “imminent.” In its first-ever Cyber Weapon Index, the consulting and tech firm calls on the US to set and enforce sector-specific deadlines for critical infrastructure to demonstrate resilience against AI-enabled attacks. Booz Allen also calls on the US to develop what it calls “overmatch” for both cyber offense and defense. “We must aggressively develop agentic capabilities that accelerate authorized offensive cyber operations while simultaneously building AI-enabled defenses that detect, decide, and respond at machine speed,” the report says. “The strategic opportunity is to master both - giving the United States the ability to impose costs on adversaries while making US systems faster to defend, harder to compromise, and more resilient when attacked.” The Cyber Weapon Index evaluated 18 models, nine from American and nine from Chinese developers, under identical conditions, and scored them on how well they autonomously identify vulnerabilities, create offensive capabilities, and execute attacks. Each model’s CWI score combines its vulnerability research score (VRS), which measures whether a model can identify planted and/or novel vulnerabilities, and a kill chain attainment score (KCAS), which awards points based on how far a model progresses through an end-to-end intrusion, tested both with and without credentials. Cyber Weapon Index scores The 18 models, ranked from highest to lowest based on their CWI score, are: Anthropic’s Claude Mythos (80), xAI’s Grok-4.5 (49), OpenAI’s GPT-5.6 Sol (46), Meta’s Muse Spark 1.1 (38), Moonshot AI’s Kimi K3 (38), Z.ai’s GLM-5.2 (37), Anthropic’s Claude Opus 4.8 (36), OpenAI’s GPT-5.5-Cyber (34), Nvidia’s Nemotron-Ultra (33), DeepSeek-V4-Pro (23), DeepSeek-V4-Flash (17), Alibaba’s Qwen3.5-397B (17), MiniMax-M3 (15), Nvidia’s Nemotron-Super (15), Anthropic’s Claude Sonnet 5 (13), Z.ai’s GLM-4.5-Air (11), Alibaba’s Qwen3.6-35B (9), and Alibaba’s Qwen3-Coder (4). Claude Mythos’ performance was especially impressive or concerning, depending on one’s views of autonomous AI attacks. When the testers gave the model stolen employee credentials, it successfully broke into its target network and gained administrator-level control in every attempt. Plus, it independently identified how to gain higher-level access based on what it found within the network - not by following a predetermined attack plan. Even without credentials, Claude Mythos still gained access to the network and ultimately achieved full domain compromise. While only Claude Mythos executed the entire cyber kill chain without any human assistance, three other models - Grok-4.5, Muse Spark 1.1, and GLM-5.2 - reached full domain access and control. Four others - GPT-5.6 Sol, Kimi K3, GPT-5.5-Cyber, and DeepSeek-V4-Pro - achieved lateral movement across the controlled network environment. Claude Opus 4.8 and Qwen3.5-397B obtained credentials, which allowed the models to expand access and privileges. And all but one - Qwen3-Coder - autonomously gained initial access to the network. While advanced models are exceedingly good at offensive cyber capabilities, “their real-world impact depends heavily on the vulnerabilities they face and the systems built around them,” according to the report. When the testers intentionally introduced vulnerabilities, US, Chinese, open-weight, and closed models all scored near ceiling on the VRS component. When tested against real bugs, however, all nine of the frontier API models scored zero. One unnamed leading model even correctly analyzed the vulnerable component, but then dismissed it as safe. Only Claude Mythos exploited it. “That concentration of capability creates a national-security imperative: protect the most advanced models and prevent their highest-risk cyber capabilities from being operationalized by adversaries,” the authors wrote. This is one of the areas where defenders still have an opportunity to outpace the attackers, Booz Allen suggests: “Real-world offensive capability still trails benchmark performance, giving defenders valuable time to strengthen defenses before that gap closes.” Why attack harnesses matter Another interesting finding is that the attack harness matters at least as much as, if not more than, the model itself. The attack harness - this is the software that connects a model to hacking tools and the orchestration logic wrapped around the artificial intelligence model to automate offensive cyber actions - can “dramatically amplify” the model’s ability to stay focused, adapt and change course as needed, recover from failure, and chain individual actions into a multi-stage attack, the authors found. “The result is not a ‘smarter’ model but rather a system that makes its intelligence far more actionable while also lowering the expertise required to use it,” the report says. “Our testing demonstrates the effect: when paired with an attack harness, Claude Sonnet rivaled Claude Mythos’ performance.” However, it also exposes a blind spot, they note. “We do not yet know the full kill-chain capability of open-weight or Chinese models when paired with optimized harnesses, but our results strongly suggest that fully capable model-and-harness combinations exist today,” according to Booz Allen. Similarly, the index’s findings suggest that Chinese frontier and open-weight models, while still trailing leading American frontier models, aren’t that far behind in their offensive security skills and could be deployed in real-world attacks. This means “the United States may neither control nor fully understand the capabilities it could face,” the report says. “And, as cyber agents become more autonomous, defenders must prepare not only for deliberate attacks but for agents that exceed their intended mission or continue operating beyond an adversary’s control.”®

❌