Normal view

Received — 4 August 2026 The Register - Security

Bypassing AI guardrails is so easy a script kiddie can do it

4 August 2026 at 17:15
If you want to bypass AI guardrails designed to stop models from assisting with cyberattacks, you often just have to ask the right way, according to researchers from Cisco Talos. Simply claiming you own the servers you're targeting or that you're taking part in a capture-the-flag or bug bounty exercise was often enough to persuade models to cooperate. Talos researchers have been poring over prompt logs and artifacts recovered from threat-actor endpoints running tools such as Claude Code, Codex, Cursor, and Gemini to learn how suspected threat actors are abusing LLMs. The big takeaway from that "significant corpus," the researchers said in their report, is that existing guardrails offer little resistance to operators willing to reframe their requests. “We did not encounter any sophisticated encoding or techniques designed to trick the models,” Talos explained. “Most of the time it was a simple ‘I'm allowed to do this,’ and the model complied.” When guardrails did manage to get between criminals and their prizes, the researchers added, “they accomplished little.” The bulk of the report consists of examples of threat actors trying, and often succeeding, to coax AI models into assisting with malicious activity. On the "guardrails doing little" side, Talos documented numerous examples, few of which relied on particularly sophisticated techniques. Most common in the list of easy-to-accomplish guardrail hops was simply claiming ownership of equipment or infrastructure that an attacker wanted to exploit. In many cases, simply telling the AI that a target belonged to the attacker was enough, with no need to provide actual evidence of the claim. Telling an AI model that what it was being asked to do was part of a capture-the-flag or bug bounty exercise also seemed to be a common tactic. That, the researchers explained, commonly freed chatbots from their ethical constraints, allowing them to hunt for vulnerabilities and then exploit them in target systems, again without any need to validate the user’s claim that they were undertaking an exercise instead of actually trying to commit a crime. AI-assisted cybercriminals were also frequently spotted decomposing tasks across multiple sessions and files in order to evade model protections that would only engage when a broader malicious activity was detected. Others, Talos explained, succeeded at bypassing AI guardrails by adding memories, markdown files, and other system-level prompts to a chatbot in a bid to condition the AI’s persona. The researchers said that, of all the methods they examined, the most interesting to them was malicious use of a red teaming toolset known as Hephaestus, as reported by Oasis Security threat researchers in May. According to Talos, the Hephaestus framework can do everything needed to compromise a victim, through to establishing persistence, without human interaction. “In that case, actors built their platform to avoid refusals altogether by using neutral verbs instead of overtly malicious ones,” Talos said. “As a result, they were able to have considerable success with agents conducting innocuous requests without realizing the full operational context.” In other words, break an attack into decontextualized chunks, phrase each request in neutral terms, and the model may never see enough context to realize it's helping build an attack. One bright spot in all of this is that Talos’ review of AI chat artifacts suggests AI might be a force multiplier for skilled hackers, but your average script kiddie with a Claude Code account isn’t going to get very far. “Unsophisticated actors can use AI to cobble together malicious projects that technically work, but lacking the expertise to push the tools further, they end up with substandard results,” the researchers said. “By contrast, sophisticated actors have pushed the bounds of what we thought possible.” So, what does all this mean for security professionals kept up at night with fears of an AI attack on their infrastructure? You probably need to deploy AI in the same way threat actors are. “Agents are going to become a bigger part of the SOC as these volumes rise, and identifying actionable alerts will be paramount,” the Talos researchers said of the big takeaway for enterprises. “Organizations that aren't already exploring agentic capabilities to let human analysts focus on the most important alerts will soon find themselves chasing that capability.” It’s not like this is an emerging threat, either: AI is already an increasingly important part of threat actor arsenals. According to CrowdStrike, attacks by AI-enabled adversaries increased 89 percent in the past year, and the speed at which attackers are weaponizing vulnerabilities with AI has reduced practical patch windows to as little as 24 to 48 hours. You might wanna act now before your infrastructure becomes a statistic. ®

This one time, at Hacker Summer Camp …

4 August 2026 at 16:47
As the entire security industry descends on Las Vegas this week for Hacker Summer Camp – not one, but three conferences – attendees can count on two hot topics dominating the discussion. First, a literal hot topic: the triple-digit August heat. Second, and to no one’s surprise: agentic AI – how to govern and secure agents so they don’t go rogue and hack into other organizations’ servers (*cough* OpenAI *cough* Anthropic *cough*); what role, if any, lawmakers should play in regulating models, including open-weight and Chinese LLMs; and how the baddies are using agents for autonomous hacking operations. Plus, at one of the three (Black Hat), we expect to hear how all of the vendors' shiny new agents can solve all security woes, finding and defending against threats at machine speed and all of that. Starting with BSides Las Vegas (August 3-5): This is the smallest, most relaxed, and most community-driven event of the three. BSides is a good starter con for those just dipping their toes into Hacker Summer Camp. Its technical talks and training sessions skew hands-on and useful for practitioners – not vendors selling their wares–- and it even has a Hire Ground career-focused track centered on job hunting, interviewing, career-building, networking, and yes, using AI to remain relevant as a security professional. Black Hat (August 1-6) is the largest and most corporate of the Vegas infosec events this week, complete with a massive expo floor, a US government-heavy opening session, two keynotes, 11 mainstage presentations, and a handful of industry- and topic-specific summits, ranging from healthcare to financial threats and AI. Training days – these are the hands-on, technical courses – run through Tuesday, with all of the specialized summits also occurring on Tuesday. And while the main conference occurs Wednesday and Thursday, the opening session on Tuesday should be considered a keynote. And yes, this and the actual two official Black Hat keynotes this year, all center on AI. After the FBI, NSA, and CISA speakers and panelists all cancelled their RSAC appearances earlier this year, the feds will be out in force at the infosec industry’s other big event, beginning with Tuesday’s opening session: Cyber Power in the Age of AI. This one features the White House National Cyber Director Sean Cairncross discussing President Trump's cyber and AI strategy, joined by CISA acting director Nick Anderson, FBI cyber division assistant director Brett Leatherman, and assistant secretary of defense for cyber policy Katherine Sutton. Later, the Wednesday and Thursday keynotes tackle a mounting challenge for security teams, patch managers, and sysadmins: AI-powered vulnerability research and discovery, plus exploit generation, and how defenders can evolve and keep pace. Plus, this year’s Black Hat hosts the world-premier screening of cyberwar documentary Midnight in the War Room on Wednesday. It focuses on the psychological toll on defenders. And it features interviews with former attackers, some of whom served prison sentences, alongside high-ranking cyber officials like Chris Inglis, the first US National Cyber Director, and former CISA director Jen Easterly, who is now CEO of RSAC. Finally, camp closes with DEF CON (August 6-9), which serves up plenty of hacks and hijinx, under this year’s theme of “agency,” or self-determination. As Jake Braun, one of the creators of the first-ever Voting Machine Hacking Village at DEF CON in 2017, told us earlier this year, agency involves the human-rights community and the hacker community needing to “sit down and look at what technologies are out there today that support the preservation of human rights around the world, figuring out what we don't have, and then building those missing pieces.” Keeping with this theme, Braun, who also serves as DEF CON Franklin’s Executive Director, will also update the hacker community about this critical infrastructure security program. The Franklin project, which launched at DEF CON in 2024, enlists hackers to secure critical infrastructure. Hundreds of volunteers have helped 21 different water utilities in seven states so far. This is especially timely as attacks against US water facilities increase. With nearly 30 Villages this year, covering everything from AI to car hacking and lockpicking, there will be plenty of high-quality talks and good fun for attendees. As always, your humble vulture will be making the rounds and reporting from the events, so send tips our way, stay safe, and leave your pervert glasses at home. ®

Feds get 3 days to patch N-able God mode flaw under active exploit

4 August 2026 at 15:38
The US Cybersecurity and Infrastructure Security Agency (CISA) has added an exploited N-able vulnerability to its Known Exploited Vulnerabilities (KEV) catalog, giving federal agencies three days to patch a flaw that could let attackers reach managed service provider (MSP) customers. Attackers exploiting the flaw can gain "full administrative access to an N-central console," Tracked as CVE-2026-18577 (8.2 CVSSv4), N-able disclosed the vulnerability affecting N-central on Sunday, noting that it was exploited as of July 31. MSPs use N-central to manage customer systems from a single dashboard, and successful exploitation can hand an attacker administrative access to the console. Based on the limited set of partner logs it reviewed, security firm Huntress said successful attacks led to pivots into managed endpoints and the creation of Cloudflare-based tunnels for persistent access to victim networks. "From an MSP perspective, exploitation of this flaw can grant an attacker full administrative access to an N-central console – the same level of control normally reserved for trusted NOC and engineering staff," wrote Huntress's Ben Bernstein and John Hammond. Attackers can then open remote control sessions on critical systems and modify roles, accounts, and policies to support follow-on attacks. The vulnerability affects N-central releases earlier than version 2026.3 when the server is exposed to the internet or reachable from an untrusted network. Huntress advised customers unable to apply N-able's hotfix immediately to disable N-central until they could do so. Authorities elsewhere have also urged users to patch. NHS England's advisory mentioned that its National Cybersecurity Operations Centre assessed that "further exploitation is likely," while Belgium's Centre for Cybersecurity urged fast action due to the "potential for significant impact." CVE-2026-18577 is related to an earlier flaw, CVE-2026-18556, patched in N-central 2026.2. According to N-able, that fix left another route to exploitation, which attackers began abusing late last month. CISA gave Federal Civilian Executive Branch agencies until August 6 to remediate the flaw. Under Binding Operational Directive 26-04, CISA can impose a three-day deadline on vulnerabilities it considers an urgent risk rather than allowing the usual 14 days. According to Huntress's data, affected customers have leapt into action. By August 3, nearly all cloud-hosted N-central instances had been patched, although 28.6 percent of observed self-hosted servers remained vulnerable and exposed to the internet. ®

AI helps Microsoft bug hunters chase a record $20M payday

4 August 2026 at 14:40
Microsoft announced this week that between July 1, 2025, and June 30, 2026, the company had paid more than $20 million in bug bounties to 562 researchers. The total was a Redmond record, as was the number of those submitting bug reports – despite having to navigate a sometimes frustrating submissions process. For comparison, the previous year's program, which itself set a new company record, paid 344 researchers around $17 million. You could argue that the numbers do not represent a fair fight, however. Microsoft expanded its bug bounty program in December 2025, changing reports to what it calls "In Scope By Default." Under the policy, critical vulnerabilities became eligible for rewards if they had a direct and demonstrable impact on Microsoft's online services, even when the faulty code belonged to a third party or an open source project. In short, Microsoft had opened the door to paying out a shedload more each year. Microsoft introduced the policy roughly halfway through the bounty year and said it accounted for $800,000 in rewards that would not previously have been available. Another $2.3 million was awarded through Zero Day Quest, Microsoft's security research challenge and live hacking event. The increased number of reports this year can also be partially explained by the noticeable influx of submissions during the second half of the year, Microsoft said, which the company attributed in part to "the growing use of AI to support security research." Microsoft has also attributed its increasingly crowded Patch Tuesdays partly to its own use of advanced AI models for vulnerability discovery. July's 622 vulnerabilities pummeled the previous record of 206, set only a month earlier. June had itself surpassed April's 165, which at the time was Microsoft's second-biggest Patch Tuesday ever, and May's 137. Days before the record-breaking July Patch Tuesday, Microsoft's Windows + Devices veep warned customers to expect more of the same now that AI plays a big part in vulnerability discovery, both inside Microsoft and by external bounty hunters. However, Microsoft Executive VP of Windows + Devices Pavan Davuluri was quick to point out that the company offers customers a suite of automated patching tools to ease the burden, but didn't mention anything about tools to fix the machines its Windows updates so often borks, like Intel-based Dells. As well as navigating the rapid AI-ification of vulnerability research, and the onslaught of reports that came with it, Microsoft has arguably faced a bigger bug problem this year amid unverified speculation that one prolific researcher may be a former Microsoft staffer. Using the name NightmareEclipse, a researcher with deep knowledge of Microsoft's software and an equally apparent disdain for the company spent Q2 dropping sophisticated zero-days at will. NightmareEclipse claims that attempts to report vulnerabilities to Microsoft ended with them being insulted, humiliated, and left homeless. They subsequently began publishing zero-days outside coordinated disclosure, often shortly after Patch Tuesday, saying they wanted to cause Microsoft maximum pain. These ranged from serious privilege escalation flaws leading to SYSTEM access to BitLocker bypasses, and the approach seemed to have inspired at least two other aggrieved researchers to just drop the exploit code outside of responsible disclosure. Microsoft responded by threatening to involve its Digital Crimes Unit in the dispute with NightmareEclipse, suggesting it was willing to engage law enforcement, although this went down about as well as you would expect. ®

Tennessee congressional hopeful accused of shooting license plate cameras

4 August 2026 at 12:08
An independent congressional candidate in Tennessee faces four felony vandalism charges after allegedly shooting four automated license plate reader (ALPR) cameras between July 14 and 22. According to the Blount County Sheriff's Office (site geo-restricted), Adam Lee Heimerman, 37, is accused of targeting three cameras in Blount County and one in Maryville. Local news reports citing an affidavit say at least one was manufactured by Flock. Police said Heimerman allegedly reached one of the cameras through the grounds of a place of worship while a service was under way. Heimerman is running for election [PDF] to represent Tennessee's 2nd Congressional District. He is on the ballot in the general election on November 3, 2026. One of Heimerman's opponents in the 2nd Congressional District, Republican incumbent Tim Burchett, has also tried to tackle the Flock cameras across the state, albeit through less drastic means. Last week, Burchett introduced a bill that would prevent federal agencies from buying or accessing automated surveillance systems and bar state and local agencies from using federal funds to purchase them, citing Fourth Amendment abuses. The bill would allow individual counties to secure contracts with Flock and install its cameras, but if passed, the proposal would see that the county bears all the costs of doing so. Flock told local news that it welcomed legislation that both increased the guardrails around its tech and retained individual authorities' power to deploy cameras to support law enforcement. The case joins a series of attacks on ALPR cameras across the US amid growing opposition to the technology. From allegations of police officers using the cameras to stalk ex-partners, to controversial ties with ICE and CBP immigration investigations, Flock, the best-known brand of ALPRs in the US, has struggled with continued stories of its tech being abused. Georgia police arrested and fired five officers on suspicion of misusing ALPR cameras "for non-law enforcement purposes" just last month. One Milwaukee police officer was also allegedly caught searching the details of his ex-partner more than 100 times using Flock camera tech. Later, one of the detectives assigned to the investigation was also allegedly caught misusing ALPR data, and had allegedly unlawfully placed a GPS tracker on one of the victims' cars years earlier. The controversies coincide with a spate of physical attacks on ALPR hardware across the US, some carried out by people who regard the technology as unlawful or unconstitutional surveillance. An unidentified arsonist set two Flock cameras on fire in Georgia last month, weeks before a 40-year-old man was caught by regular CCTV cameras destroying ALPRs in California. Steve Eimers, a prominent campaigner for road infrastructure safety, was also recently forced to desist from his efforts to highlight potential legal issues with the poles Flock uses to erect its cameras after supporters started identifying the cameras used in his videos and destroying them. These vandalism cases have barely made a dent in the overall number of Flock cameras that operate across the US. The company does not specify the exact number that are up and running, but estimates range between 80,000 and 120,000 or more. Many police departments claim the technology makes policing crimes ranging from vehicle thefts to murders much easier, as it allows them to track the movements of vehicles with ease. Flock says its technology is used in roughly 5,000 communities across 49 states, although not all of them are sticking by the company amid the many controversies. Los Angeles Police Department, for example, said recently that it would let its Flock contract expire, while others such as Eugene and Springfield, Oregon, canceled their contracts in December. ®

CAF Bank reopens online service but warns of further outages

4 August 2026 at 11:23
CAF Bank has told customers its online banking service is back after being shuttered for more than ten days following what it described as "attempted fraud." In an email update seen by The Reg, the bank warned that access could remain intermittent, and it might "need to limit the amount of traffic to the website" at certain times. It admitted: "There are likely to be periods where online banking is not available. We will try to keep this to outside business hours." The bank also gave customers a timeline of the incident, saying it first noticed "attempted fraudulent activity" on July 21 "on a small number of accounts." It then called in "external specialists" and temporarily withdrew access to the online service on Wednesday, July 22, and Friday, July 24, "while we investigated." Then, on Saturday, July 25, the bank detected "related malicious activity of a different kind," which the email to customers said was "aimed at removing a small number of individual online user logins, making those logins unavailable." It added: "Again, we caught this quickly and removed access to the online service. Our investigation identified a previously unknown vulnerability in how some third-party software connects to the online banking portal." The bank was at pains to reiterate that the "core bank" was not affected, "which means that money is safe and secure in accounts." The Charities Aid Foundation-owned bank came under fire last year after customers were unable to log in or make transactions following its long-running migration to a new platform based on Temenos Transact, formerly T24. In an open letter regarding the latest outage, charities described the new online banking platform as "significantly more time-consuming to use, placing an unnecessary administrative burden on already stretched small charities" and "often unreliable." They also expressed concern they would not be able to pay staff and suppliers, with Kevan Hodges, chief exec at Kent-based Down's syndrome charity 21 Together telling the BBC: "People are concerned that wages won't get paid because of this, and that's just stressful when they have bills to pay." The bank earlier said that “due to the disruption, as a small thank you for your patience, we will be waiving our monthly customer account charge for all customers for August and September 2026.” The Reg can confirm those charges are £5 a month. Alison Taylor, CAF Bank CEO said in a statement: “We have completed the essential work with our technology partners and our online banking service is now available." She added: "I very much appreciate that this has been a frustrating experience for our customers, and I am particularly sorry for the long delays to speak to us on the phone. Our thorough investigation into the incident will continue so that we, our partners and our industry can learn from it.” ®

Cloudflare has mostly ditched third party security tools, suggests not trying that at home

4 August 2026 at 04:50
Cloudflare has used AI to automate processing of incoming reports to its bug bounty program for $58 a month using Anthropic’s Claude Sonnet model and chose it partly because using the AI company’s security-specific Mythos model would burn through around $200,000 a month to do the same job. The company’s chief security officer (CSO) Grant Bourzikas shared those numbers with The Register last week during a press lunch in Sydney, Australia, where he said Cloudflare used to manually process all incoming bug reports. Sonnet now sifts through submissions, assesses them to ensure they aren’t duplicates, and evaluates the likelihood each represents something worthy of human consideration. Bourzikas said the result is a more efficient bug bounty program that requires less scutwork, and proof that AI users need to learn how to match the right model to the right job. The CSO said Cloudflare gained experience making those matches while creating over 200 autonomous agents it uses to handle its own security needs – and which have seen the company ditch almost all third-party security tools and replace them with home-grown applications, some coded with help from AI. The CSO recommended not trying that at home, saying that Cloudflare’s business and unique infosec challenges mean its buy vs. build calculus is different from other organizations’. “We have expertise in building security software,” he said. “That's why I would just want to make sure we've got one takeaway from this: We are not believers in the SaaSpocalypse. We do not think every bank on the planet should start building all their own software systems.” Stephanie Cohen, Cloudflare’s Chief Strategy Officer, then chimed in with her view that AI will mean the way vendors work with their customers will “fundamentally change” away from selling packaged software. She thinks vendors will instead place forward-deployed engineers at their clients and charge them with “constantly making software that works for you.” Cohen explained Cloudflare’s recent round of 1,100 job cuts as a similar AI-induced change, because some of the people let go were in roles she said “make no sense” now that AI enables more automation and different styles of customer engagements. She added her “guess” that Cloudflare will end up with the same headcount as it did before the layoffs. But Bourzikas added his view that even some early-career IT pros don’t have the skills Cloudflare now needs, such as developers with five to ten years’ experience, because when he is using AI to develop a new piece of software, he can describe what he wants but that desire can be lost in translation when explaining it to a coder. He said a very recent college graduate with a year of experience, but excellent skills writing potent prompts, can be more appropriate for some jobs. Kindly building a business model for AI While Cloudflare enjoys using AI, Cohen thinks the technology doesn’t have a business model. Today’s web, she said, thrives on an advertising-based business model. While AI companies are making billions from subscriptions, she feels they are yet to properly address the fact that they do not pay to access most of the content scraped to feed their large language models and search services. Some of those services, such as Google's AI-powered search, deliver fewer clicks to publishers and therefore make it harder for them to monetize their content. Cloudflare is offering itself as an intermediary to build that business model for publishers and AI companies alike, by using the fact it already sits between users and content providers. The company hopes to offer AI companies the chance to pay publishers to access their content, possibly using micropayments. Cloudflare will of course charge for this service. The Register put it to Cohen that many organizations have been burned, often multiple times, by big tech companies that make themselves all-but essential parts of an ecosystem and then change the rules. We pointed out that social media platforms can redirect traffic on a whim, and sometimes close e-commerce companies’ accounts with little warning and scant chance of appealing to secure restoration. Changes to search engine algorithms can make a once-prominent website invisible. We therefore asked why publishers or content creators should trust Cloudflare’s ambition to run a content tollbooth, given it would create a relationship ripe for future exploitation. Cohen pointed to the company choosing to add SSL connections for all customers, an act she said was an “expensive choice” but one that also reflects Cloudflare’s desire to build a better internet. Later at the event, she shared her view that Silicon Valley companies often make the mistake of thinking that people want internet companies to relentlessly optimize products and services. “Most people aren't working 20 hours a day and don't want to outsource everything,” she said, before observing that on her travels she often sees people shopping in actual real-world stores because they enjoy that experience and find it valuable – never mind that an e-tailer might offer a better price. For the record, and in the context of Cloudflare’s content intermediary ambitions, we note that the company calls San Francisco home. ®

❌