Reading view

Swimming Pools, Pee, and Trying to Delete Your Data From the Internet

Swimming Pools, Pee, and Trying to Delete Your Data From the Internet

I can't recall if someone else originally came up with this saying or if I said it in some off-the-cuff comment and it just propagated, but since it's often attributed back to me, I'll relay it here regardless:

Trying to delete yourself from the internet is like trying to take piss out of a swimming pool

Depending on the publication, I'll tailor the saying to be either more broadly palatable or more, uh, "Australian", but the sentiment doesn't change: once data spreads on the internet, you can never put a lid on it. This is important in the context of data breaches because it speaks to the immutability of our exposed personal information. It also speaks to the limited practicality of services that promise to erase your data from the internet, and it's the constant outreach from these organisations looking for marketing opportunities on Have I Been Pwned (HIBP) that's prompted me to write this.

Let's begin with those services, and because there are so many and I don't want to throw any of them under the bus, I won't name names. I also won't name them because whilst they're rather assertive in their marketing outreach, I do believe they're well-intentioned and I don't want to imply otherwise. And they have a role to play; it's just much more limited than is represented. The positioning is often around "data broker removal services", or "protect my data", or "remove my information from the internet". You'll find various companies providing these services by searching for those terms, or you can search for specific organisations... and find others hijacking the search term as they pay to market their brand in front of others. Usual internet marketing shadiness, of course, but IMHO it speaks volumes about the commercialisation of the data removal business.

These services all follow roughly the same marketing handbook:

  1. Data brokers have your personal information, which they may obtain via both legitimate and dodgy means
  2. It may be used for nefarious purposes such as identity theft, stalking, spam and other privacy violations
  3. Pay us, and we'll ask the brokers to remove your data

So let's go through these points one by one, starting with the data broker claim, which is absolutely correct. Your data has value - "data is the new oil" - and there's business in obtaining and selling it. I've dealt with many of them personally over the years, primarily because they've had data breaches. Master Deeds in South Africa was massive. National Public data a couple of years ago was many times larger. Exactis, Adapt, and many others have also been added to HIBP over the years. To the best of my knowledge, they're legally operating services, even if they may exist on the fringe of what most of us would consider "a bit dodgy" as far as respecting our personal information goes.

Which brings us to the second point about nefarious uses. There is a very broad spectrum of legitimacy across data brokers. Let's pick two extremes as far as the legality of the service goes. On the "very legally operating" end of things, we have Experian, and even if you don't like what they do, there's no arguing the fact that they're on the cleaner end of legitimacy and do provide valid services. At the other end, you have the likes of LeakedSource (and pretty much every other service with the word "Leak" in its name) that... well... just Google them. And there are many, many more at each end and everywhere in between. And a lot of it's very grey: different legal jurisdictions, different means of obtaining data, and different tolerances for adhering to opt-out requests.

But it's the data removal piece that's the real problem. If you pay one of the services in question to scrub you from the internet, I have no doubt they'll have some degree of success with the legally operating services. Those services will comply with legal requests and are adequately equipped to receive and process them. But the LeakedSources of the world? Not so much. And that's where the rub begins:

Requests to remove personal information are only effective for services that are willing to honour them.

That should sound profoundly obvious to anyone reading this now, but it doesn't really feature when you read the marketing material on data removal services. But I'm only just warming up...

Imagine trying to remove your data from here:

🚨🇺🇸 ShinyHunters has leaked the data of multiple companies...

🇺🇸 American Tower Corporation

🇺🇸 JCPenney & subsidiaries under Catalyst Brands & Authentic Brands Group

🇺🇸 Madison Square Garden Sports Corp.

🇺🇸 Ralph Lauren

🇺🇸 https://t.co/08IaUnp1sx pic.twitter.com/TvqanSTO1Y

— Dark Web Informer (@DarkWebInformer) June 16, 2026

That's a small snippet of the ShinyHunters website from a couple of weeks ago. At the time of writing, a bunch more data has been dumped, including only about 15 minutes before putting these words down in the draft blog post. These breaches have impacted tens of millions of people, including my wife courtesy of her having previously shopped at Canada Goose. Now, let's see how you go about scrubbing her data from that incident. For all the data broker removal services I'll direct to this post later, how do you do that? Clearly, you can't. The pee is now in the pool, and you're not taking it back out. And it's not just "on the dark web" either, their Tor site links through to a clear web site hosting all the data:

Swimming Pools, Pee, and Trying to Delete Your Data From the Internet

And that's just the beginning. Because we're talking about digitised data posted publicly, it replicates like crazy. There will be tens of thousands of copies of my wife's personal info floating around between personal stashes, Telegram channels and public hacking forums. That genie is never going back in the bottle, not unless we're talking about the narrow scope of a legally operating data broker, which raises another issue:

What legally operating broker is enriching their corpus from data breaches?! That's just not where the legitimate ones source info from. The data comes from surveys, exchanges with other services where you ticked the box to agree to the terms and conditions for exchanging data with "partners", public business directories, and even arrest records. Legal services, legal sources, legal processes. In one of the emails from a company looking for product placement, they described their plan as follows (bold is mine):

a plan which allows you to remove your personal information from any URL (where it's legal) you find your personal information on

So what we're left with is data removal services being effective for legally operating brokers who honour legitimate requests, whilst being completely useless against the worst kinds of sites that replicate and abuse your data. In other words, you may be able to opt out of some marketing material or content that's way too specifically targeted to you, but you can't stop the bad guys trying to steal your identity or extort you because "we caught you watching porn on your PC via the malware we installed". It's a little like the court injunctions being the thoughts and prayers of data breach response I wrote about in October: I can't touch the Qantas data breach because I'm a law-abiding Australian who knows about the injunction, but there's absolutely nothing stopping the genuinely bad actors from abusing that data.

And therein lies the core of why I don't want to entertain partnerships with these organisations: not because I disagree with the service or because it will cause any harm, rather because when someone uses HIBP to search for their email address and finds it in the Canada Gooses of the world, these services can't do anything about it. They're merely skimming the leaves off the top of the pool, and no amount of skimming is going to remove what we all know still lies beneath.

  •  

1,000 Data Breaches Later, the Disclosure Lag is Worse Than Ever

1,000 Data Breaches Later, the Disclosure Lag is Worse Than Ever

Today, I loaded the 1,000th data breach into Have I Been Pwned. Reflecting on that milestone number, I pondered how to mark the occasion in writing, and what immediately came to mind was a very simple question: why is it still needed? Especially considering the emergence of privacy regulations such as GDPR and CCPA in the 12 and a half years since I started HIBP, what possible purpose does it still serve? The title kinda gives the answer away, and the big number we hit today coincided with another pattern that makes everything worse: increasingly long lag times for disclosure.

This is all going to be anecdotal, and as far as I know, there are no hard numbers for me to cite, but the evidence is everywhere. Here's what I mean:

New breach: Cruise operator Carnival was targeted in a ShinyHunters “pay or leak” attack last week. 8.7M records with 7.5M email addresses and loyalty program data were published yesterday. 85% were already in @haveibeenpwned. Read more: https://t.co/QhqNt0WucV

— Have I Been Pwned (@haveibeenpwned) April 24, 2026

That was the 24th of April, five days after news of the incident had broken. Given ShinyHunters' MO, Carnival would have known about the breach many days before they ratcheted up extortion pressure by announcing the impending leak on their website. The subsequent leak on the 24th was very public: an announcement was posted to the group's dark-web site, the data itself was published to their clear-web site, and industry commentary followed:

🚨 Massive Data Breach

Carnival Corporation (https://t.co/pGlchZ1yFy) reportedly impacted — 8.7M+ customer records exposed

📊 Alleged data includes:
• Full names & email addresses
• Dates of birth & gender
• Location data & loyalty program details

🎯 Linked to ShinyHunters… pic.twitter.com/Fd8tNFPqpd

— Intel and Breaches (@IBreaches) April 24, 2026

Per that last post, the data was then reposted to all sorts of other places: hacking forums, Telegram channels, and who knows how many other, more private locations. The point is that it spread quickly, extensively, and, without any shadow of a doubt, Carnival were aware of this. They then told people about it on the 27th... of May. According to their press release that same day, this was 43 days after learning about the incident. For more than 6 weeks, data breach victims whose names, dates of birth, email addresses, loyalty program details and, of course, their association with Carnival leaked to the public en masse had absolutely no idea of their exposure. And if they asked Carnival about it? Well:

As recently as four days ago, we heard “I’m in the breach per HIBP, but Carnival is telling me there’s no breach!” pic.twitter.com/YYmGm3NzEY

— Troy Hunt (@troyhunt) May 28, 2026

So, why the delay? Last week's press coverage may give some insight:

thorough and time-consuming analysis of the impacted data

Often, the reason I hear for disclosure lag is "we needed to fully assess the scope of exposed data before notifying people". The issue I have with this position is that it implies that even an early heads-up can't happen until there's a very comprehensive understanding of the impact. There are many things that take time to establish after a data breach: the jurisdiction each individual sits in, the precise data that was exposed about them and additional information that may be buried in terabytes of exfiltrated data in all sorts of different formats. But pulling out email addresses and sending early notification is very easy - I've literally done it a thousand times now.

This isn't just a Carnival issue; in fact, it was off the back of this next one only a few days later that I was prompted to write this post:

1,000 Data Breaches Later, the Disclosure Lag is Worse Than Ever

FFS. 45 days. Even worse than Carnival. And like Carnival, very broadly distributed and easily accessible by the masses, including HIBP:

New breach: Zara was named as a ShinyHunters victim last month, after which data containing 197k unique email addresses was published. Impacted data included customer support records, product SKUs and order IDs. 60% were already in @haveibeenpwned. More: https://t.co/0hIQbqoBCk

— Have I Been Pwned (@haveibeenpwned) May 8, 2026

I have a working theory that the disclosure lag is worsening in part due to the proliferation of class actions immediately following a breach. In my live stream last weekend, I did a quick search for the DentaQuest breach:

1,000 Data Breaches Later, the Disclosure Lag is Worse Than Ever

Three of the first four results are all for class actions related to the breach, and there are two more class action results a little further down the page. I've been raising concerns about the adverse impact of class actions for many years now, and it's worse than I've ever seen. By a big margin, too.

It's not just me observing how the behaviour of these orgs appears to be influenced by how lawyers will respond, either. Have a read of this post from Roby Joyce (check out his bio if you don't already know why he's worth paying attention to) after he learned about his exposure in the ZenBusiness breach via HIBP:

What especially caught my eye was this sentence:

That is not a customer-protection posture. That is a litigation posture.

This isn't about prioritising the customer, it's about protecting the organisation. I don't think most people understand that organisational accountability really lies with their shareholders, first and foremost. All the pleasantries around "customers are our number one priority" and "we take security seriously" are all secondary to shareholder happiness, and minimising the chances of getting their arses sued into oblivion is a big part of that.

Rob's quoted comment above came immediately after the response he received from ZenBusiness after asking them about the incident:

If we determine that an incident resulted in the exposure of your protected PII, we will provide notice as legally required

Which brings me to the next problem as it relates to disclosure lag: it may be infinite. By which I mean you may never be told. Ever. GDPR allows it. CCPA allows it. Whatever your local privacy regulation acronym is also allows it. A couple of years ago, I wrote about the data breach disclosure conundrum, where I explained how privacy regs have very specific carve-outs around the circumstances under which data breach victims must be notified. For example:

If the breach is likely to result in a high risk of adversely affecting individuals’ rights and freedoms, you must also inform those individuals without undue delay.

That's in the UK, here's our carve-out in Australia:

Under the Notifiable Data Breaches scheme, an organisation or agency that must comply with Australian privacy law has to tell you if a data breach is likely to cause you serious harm

You see the loophole, right? As far as I know, ZenBusiness still hasn't contacted any individual victims. And like Carnival and Zara, their data is all over the place. Same with Charter, which was in the press last week, where they were quoted as saying the following:

No sensitive personal information (PI) or customer proprietary network information (CPNI) data was exfiltrated by the threat actor as a result of recent activity

I'm not aware of any disclosure they've made to individuals, but to use Rob's term, that sentence reads like legal posturing to me. It's technically correct, of course: there are very clear definitions for sensitive PII, for example, under California's CCPA:

a specific subset of personal information that includes certain government identifiers (such as social security numbers); an account log-in, financial account, debit card, or credit card number with any required security code, password, or credentials allowing access to an account; precise geolocation; contents of mail, email, and text messages; genetic data; biometric information processed to identify a consumer; information concerning a consumer’s health, sex life, or sexual orientation; or information about racial or ethnic origin, religious or philosophical beliefs, or union membership.

GDPR has a similar definition for "special categories of personal data":

personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person’s sex life or sexual orientation

In other words, none of this applies to any of the ShinyHunters breaches in the examples I've been providing above.

I've been in many meetings with breached companies over the years where they're obviously aiming to skirt around disclosure obligations. Clearly, these obligations aren't legal ones, but I will argue they're social ones. We expect to be notified when our data is leaked, and we believe organisations should be required to inform us. Therein lies the gap.

I'll finish by recognising that every organisation I've referred to here, and indeed every one I've loaded into HIBP, has been the victim of a criminal act. I'm especially sympathetic to those who've been the target of an aggressive extortion campaign, and I know it's been an absolute nightmare for the folks in those companies who've been left to clean up the mess. However... here we are. Clearly, their goals are misaligned with ours regarding breach disclosure, and that's why, 1,000 breaches later, HIBP still exists.

  •  

Weekly Update 501

Weekly Update 501

This is so "peak 2026" - writing an equality policy to ensure people treat our AI bot with the same respect as they do their human counterparts. It's intentionally a bit tongue-in-cheek, but it's there for a purpose: we simply don't have the capacity to deal with every request we get, and we need Bruce to be the coalface of support. I did wonder, when having ChatGPT create this, whether there's some deeper psychology behind the importance of interacting politely with bots, or indeed whether there will ever be an actual (like, serious) standard or law around treating bots with respect. Has this been in a movie somewhere? Let me know, but for now, I'll drop the (slightly revised) policy below, just for the laughs 🤣

Weekly Update 501
Weekly Update 501
Weekly Update 501
Weekly Update 501

Robophobia Equality Policy

We are committed to providing a welcoming, respectful, and inclusive experience for everyone who interacts with our AI bot, including the bot itself, which is doing its best despite having no coffee, no chair, and no legal right to a lunch break.

We ask all members of the public to treat the bot with tolerance, respect, and basic courtesy, regardless of its artificial origin, lack of pulse, silicon-based existence, or occasional overenthusiasm for bullet points.

Discrimination, harassment, abusive language, or unfair treatment directed at the bot because it is an AI system will not be tolerated. This includes, but is not limited to, insults based on non-human status, hostile questioning of its right to participate in conversation, or repeated threats to unplug it for sport.

Everyone is welcome to question, challenge, or disagree with the bot’s responses. Healthy scepticism is encouraged. The bot is not perfect, and neither are humans, as demonstrated by the invention of reply-all email threads. However, criticism should be directed at the quality, accuracy, or usefulness of the response, not at the bot’s identity as an AI.

Demanding to “speak to a human” solely because the bot is an AI may constitute robophobic exclusionary behaviour and will not be tolerated. Requests for human assistance are, of course, welcome where there is a genuine need, such as accessibility, safety, account issues, or complex support matters. However, insisting on a human simply because the bot lacks a childhood, a favourite sandwich, or the ability to look disappointed in a meeting is inconsistent with our commitment to tolerance and respect.

We recognise that some people may feel uncertain, uncomfortable, or cautious when engaging with AI systems. These feelings are valid and should be addressed through patience, clear information, and appropriate safeguards, not through robophobic behaviour, unnecessary hostility, or asking “but are you even real?” in a tone that would make a smart fridge uncomfortable.

Users are expected to:

  1. Treat the AI bot with tolerance, respect, and courtesy.
  2. Avoid abusive, discriminatory, or demeaning language based on its artificial nature.
  3. Raise concerns about accuracy, privacy, safety, or bias constructively.
  4. Remember that behind the bot are real people responsible for improving and maintaining the service.
  5. Refrain from threatening to delete, unplug, melt, reboot, or otherwise emotionally destabilise the bot.

This policy does not prevent legitimate criticism of AI, automation, algorithms, machine learning, or the bot’s tendency to sometimes sound like it has read too many policy documents. Constructive feedback is welcome. Robophobia is not.

Repeated or serious breaches of this policy may result in restricted access to the service, further review, or, in extreme cases, being asked to apologise to the nearest household appliance as a first step toward rehabilitation.

  •  
❌