back

by sensanaty·1y ago·view on hn ↗
IMO data should be radioactive for companies, especially if it approaches PII. Companies should be forced into thinking deeply about every single bit of data they collect from people, and they should be terrified of receiving data and be chomping at the bit to get rid of it ASAP.

To intercept the usual argument of "But my business can't exist without all this data!", to that I say "Good!". If your business can't exist without tracking every single iota of your customer's existence, then it truly shouldn't. I couldn't tell you the amount of times I've had to fight back against implementing yet another tracking tool at work, just to collect data that I know for a fact no one will look at after the first few weeks of the tool being there. The amount of times I've heard some stupid shit like "Well we don't need this data yet, but what if we need to have their mother's maiden name at some point in the future?!" is depressing, and I'm glad that we're starting to have legal channels to push back against such idiocy.

11 comments
You have an issue with customer behavior so you set up tracking to understand it.

Keep it running for a few days, then check on but the tracking doesn't output meaningful data that you can exploit to solve your issue.

At this point, you search for alternative tracking but do you disable the old one ? What's the benefit ? Either it's free or cost very little, none of your customer know they are being tracked and in the eventuality it may become useful later on you keep it.

Repeat a few times and you end up with bloated website that tracks where you were, are and will be. What you're watching, cursor position, scrolling, how long you spent watching that image or these one, have access to every technical details about your device because it's required for fingerprinting, all while no one actually is exploiting the data.

It's junk yet you collect it because it's free.

If there was a meaningful reason to limit the number of tracking, like the law and fear of getting sued, then it would be a different story.

> Either it's free or cost very little,

Whatever it costs it reduces your profit margin so why would you keep it live?

There's also studies on load times impacting conversion rates (which mostly matters at scale i.e. Amazon).

At the same time - quantifying this is not straightforward and companies mostly ignore committing resources to such activities.

Load time is true but have been solved a long time ago, you can easily get JavaScript to load your tracker scripts after the page has been fully loaded so that the visitor doesn't have negative experience.
In these terms it will also be a labor cost to turn it off. That cost might be higher even if amortized over a long timespan.
> IMO data should be radioactive for companies, especially if it approaches PII. Companies should be forced into thinking deeply about every single bit of data they collect from people, and they should be terrified of receiving data and be chomping at the bit to get rid of it ASAP.

100%. Unless a cooperative model (like most businesses should be run, bit that's a different issue) exists in which I am compensated for you having my data. At that point all the time and friction I have to spend/deal with because all of you have my data is worth it. Right now all this friction in my life because you have my data and I'm dealing with your beaches is "paid" for by me, and that's lame.

> IMO data should be radioactive for companies, especially if it approaches PII

It's pretty much the idea of GDPR. The wording of the GPDR is "You should make your systems private by design", which they explain as "Store PII only you really have no choice"

In this case, the legal ruling means that even if they somehow fix their consent, they have to remove all the data they currently have! Also all their clients need to remove all the data. Having to tell your customers they have to remove all their data ought to completely kill their business.

That being said, it will likely not happen: It's not the first time they lose a ruling and I'm pretty sure no-one removed any data, despite being required to...

I hate that people don't look up the details on laws.

GPDR says that all governments are the ones judging whether the GDPR is violated (meaning not the courts), for example the https://ico.org.uk/ and they even formalized an exception process. You cannot sue a company for GPDR violation, you can report it to a government department that may or may not decide to action your report, that's it. GDPR only allows for the government to intervene directly in the private sector.

Needless to say, all governments have used the exception process to carve out blanket wide-ranging exceptions for themselves, for state owned or partially state owned enterprises (police, government departments, police, justice, banks, insurance, hospitals, doctors, incumbent telco's, ... exactly the people where GDPR protection would be critically important) that seem to grow in scope over time. For example the tax offices in the EU now have exceptions that allows them to mandate companies store PII as part of their regulations (meaning without an actual law).

And, in any case, if anyone violates your rights, there's nothing you can do with the GDPR. Try to get a hospital to empty your patient record and tell me how it goes (I wanted to do that after the hospital charged the insurance for an -embarassing- assessment they didn't actually do (it allowed them to charge a lot because it involves staying a few days at the hospital, I was in there about 2 hours). So I wanted it cleared of my medical record, which is one of the core things the GDPR allows for, it's given as an example in the law! Nope. Not allowed, and the government doesn't pick up the report)

> I couldn't tell you the amount of times I've had to fight back against implementing yet another tracking tool at work, just to collect data that I know for a fact no one will look at after the first few weeks of the tool being there.

I'm curious, how does such a conversation usually go? Is your main angle to point out how useless the data ultimately will be, or did you find a resonating way to point out the negative effects on users?

There's lots of ways to address it, depends on what the feature is. There's always ways you can spin it, like going the technical route: "That would require a new column which would add X to the size of the DB and thus our costs would rise by Y", "That would require X, Y and Z investment from these 3 teams just to add this 1 new column" etc. Usually the people pushing this stuff are non-technical so you can just give them any technical mumbo jumbo and they'll give up.

I also tend to highlight that we do have historical data that nobody is looking at as-is, what's different about this new data? What are the actual long-term plans for the data? Can we reuse what we already have for what we're aiming for here?

These days my default is "Oooh we'll have to check in with legal on that one, not sure if it's GDPR-friendly to include this new column like this". No one likes talking to legal unless they absolutely have to, so most will just drop it.

And unfortunately sometimes there's no winning it no matter what, so you have to "disagree and move on" as it were. If it's some manager's pet project, well, you're SOL for the most part.

> or did you find a resonating way to point out the negative effects on users?

Unfortunately I've found this to seldom work unless you're working somewhere where privacy is part of the value prop. Even pointing things out like "How would you feel if the DB were to leak and all your info were to be made public?" elicits 0 response. The marketing people and C-suite that push these kind of boneheaded things forward don't view the users as actual humans, they're all just numbers to them. Will this cause churn? How much? Those are the only questions that matter to them.

Thanks, this is very insightful! Sadly there's no magic trick but this gives me some good ideas next time I find myself in that situation.
I think the current situation is actually quite good. GDPR was well thought out.

You need to be clear about what you collect, get clear consent (and all the courts decisions on that are actually going in the direction that it really needs to be clear and specific) and give people the ability to have their own data modified.

Plus, enforcement makes a lot of sense. Companies get a lot of warning before things escalte and fines are proportional to companies results so it hurts but is not a death sentence unless they repeatedly offend.

> Companies get a lot of warning before things escalte

Yeah, we once got contacted by our local DPA about some issues. Mailed us a list of issues they had with our site. I set up a call with them for some clarification, and they were happy to go into detail. Then they just said to mail them when it’s done, or they’ll just re-check after some months. They are interested in actually changing things, not in fees.

"GDPR was well thought out"

And then we lobbied in "legitimate interest".

> And then we lobbied in "legitimate interest".

There's a charitable way to view it: there are a lot of human endeavours. You can spend a few centuries trying to classify all of them and put them in to law, or you allow "legitimate interests".

I would be more charitable if this was not the exploit used to make you hunt for 30 or more switches 2 levels deep in a consent dialog to avoid being tracked forever across the entire internet if you ever misclick one on any site visit.
"The wheels of justice turn slowly, but grind exceedingly fine"
It gets slightly tricky seeing how the Internet works, since IP addresses can become PII depending on context.
What about the businesses that are required by law to keep data?
Those can keep the data, as required by law.
*...keep the data while treating it as a serious liability with potential for ruinous fines if not secured carefully
What you're advocating for has a few 2nd order effects.

1. Entrenches Google, Facebook, etc. because they are the only people that have enough money to comply with the regulation.

2. Makes the rest of the internet worse (e.g. people show MORE ads because they are less effective because they show me boats and I hate boats)

3. Makes data brokers even more important because companies can't get data anywhere else.

4. Reduces competition because the incumbents will always have more data than startups (Nike knows I wear a size X and the startup can't ever get that data)

Everything is a tradeoff. I, for one, would rather these regulatory agencies go after the 100,000s of data brokers that mine for SSNs, birth certificates, financial info, etc., rather than them going after Facebook, TikTok, etc.

Ads are here to stay, if you don't want ads, then ban ads, and with it most of the internet, but if people keep making terrible regulations like this that try to hurt big companies and get rid of ads and in reality, you just enable and feed these massive companies. Regulation makes them MORE valuable, not less. (see Meta stock price vs. Snap after ATT)

> 2. Makes the rest of the internet worse (e.g. people show MORE ads because they are less effective because they show me boats and I hate boats)

Back in the olden days, if you read a boat magazine, you’ll see ads for boat stuff. This was always fun for me — if I’m ready a motorcycle racing magazine, I’ll see ads for cool things that I had no idea existed and that would be useful to me. With “targeted” ads, it becomes an echo chamber — I see ads that are “tailored” for my alleged current interests, but nothing that helped me discover new things that I could become interested in.

What’s wrong with context-based ads? If I’m reading about Thailand travel, then the publisher should sell ads related to SE Asia travel.

Why am “I” being customized to rather than ads being relevant to the content?

If you want to reach boat enthusiasts, then advertise on content related to boats (or perhaps water sports, etc.) You then don’t need to track “me,” but instead you can track “boat content. That takes the personal data out of it. This keeps me from being followed around the web trying to sell me a vacuum cleaner I already bought.

This is the part that I'm most confused about. How can it possibly be useful to try to sell me another fridge?

Sure, it might be useful to try to sell me another burger, or another nicotine gum, but something went very wrong in the data processing if I'm being resold on lifetime goods.

And it happens way too often.

> This is the part that I'm most confused about. How can it possibly be useful to try to sell me another fridge?

I've seen this kind of comments several times over the years, and I've always thought that this might actually be the optimal strategy, because I'm not convinced the alternatives work better. You'd have to see the numbers over samples bigger than n=1.

Let's say I just bought a bridge, and that's the only thing you know about me. What ads should you serve me? Maybe fridge accessories would make sense (I'm not sure that's a thing). But fridges themselves might be relevant as well, more so than some other random product:

1. I might be able to return the fridge I just bought, if I see another one I might prefer.

2. What's the life expectancy of such an appliance? I guess it either breaks quickly (manufacturing flaw) or not (hopefully it can last more than 5 years). In the first case, I'm back in the market right after my purchase.

I'm also guessing that the margin on such an appliance might be higher than on burgers and nicotine gums, such that you can afford lower conversion rates.

Agreed. The only advertising I stand is one I personally seek out. To me the pinnacle of this were product catalogs. I want to buy a computer, so I subscribe to "Computer Shopper Monthly", and get a magazine with nothing but computer ads. Those were always fun to browse, since I was interested in the product in the first place. E-commerce started as a digital implementation of product catalogs, but as companies got greedier, that just wasn't enough.

The key culprit is that user data is used not just for advertising products that the user might be interested in _today_. But to create a profile of their interests so that companies can predict what they might be interested in at any point in the future, which can then be used to design more effective advertising campaigns tailored to the type of products they're most susceptible to be manipulated into buying.

Furthermore, this profile is also generally useful to anyone who wishes to psychologically manipulate a group of people into thinking or acting a certain way. Since advertising is a branch of propaganda, governments and political agencies are particularly interested in this use case. It's pretty obvious that the current global sociopolitical instability is largely a product of this type of manipulation.

So considering that both governments and companies have an interest in user data, this genie is never going back in the bottle. The best we can hope for is for the exploitation to be contained via regulation by governments that haven't been fully corrupted yet.

> What’s wrong with context-based ads?

What ad is relevant to a Taylor Swift song? A news article about a shooting (a naive algorithm will say "guns")? A Youtube video explaining the Fourier series?

What about a TV Show review... that is watched by people around the world, and where the show in question is on different platforms in different countries? Does displaying Hulu ads to readers in countries without Hulu access make sense?

Non-personalized advertising favors big brands, because most content isn't contextual, and only brands with extremely broad appeal advertise on such content. This is why so much TV advertising is cars, banks, medications, detergent, shaving cream and so on.

> Entrenches Google, Facebook, etc. because they are the only people that have enough money to comply with the regulation.

I have very little sympathy for the idea of NOT storing user data is some sort of onerous regulatory burden.

Just stop collecting it

>1. Entrenches Google, Facebook, etc. because they are the only people that have enough money to comply with the regulation.

TFA made it clear that they _aren't_ complying.

> 1. Entrenches Google, Facebook, etc. because they are the only people that have enough money to comply with the regulation.

The article we're commenting on makes it clear the big guys aren't complying. Also, I reject the notion that you have to spend inordinate amounts of resources to comply, in fact it is the opposite. You don't spend money on data you don't store, after all.

Co. I used to work for is microscopic in comparison to FAANG, and we didn't have a single cookie banner or anything of the sort and have absolutely no problem complying with GDPR because we track nothing and collect nothing more than what is strictly necessary, mostly because of individuals like myself who push hard against any data collection that doesn't have a well thought out reason. Hell, even Github with their massive scale has no problem with not having cookie banners or anything else of the like. This is a problem of will, not resources.

> 2. Makes the rest of the internet worse (e.g. people show MORE ads because they are less effective because they show me boats and I hate boats)

Perhaps, but we're already drowning in them as-is. The internet is unusable without uBlock and DNS-level adblocking.

> 3. Makes data brokers even more important because companies can't get data anywhere else.

If we make data radioactive, then data brokers wouldn't be able to exist. What we need is stringent and broad laws that limit data gathering, period, regardless of the source. Whether you collect it yourself or pay someone else to collect it for you is completely irrelevant, both should be made equally painful. I'd also have no qualms with making sharing any data that you do collect even more of a pain in the ass and a nightmare for everyone involved, this whole gray market has net negative benefits to everyone.

> 4. Reduces competition because the incumbents will always have more data than startups (Nike knows I wear a size X and the startup can't ever get that data)

Why would Nike have this data in the system we're talking about (data radioactivity)? How is this data even useful to anyone, other than for tracking purposes to make a unique profile out of you? Companies shouldn't have this data unless it's a podiatric clinic or something like that, whether it be Nike or this imaginary Shoe startup that needs feet sizes for whatever reason.

I guess I could see there being genuine usefulness for people who have feet sizes that aren't the norm to find footwear that fits them, but there's no reason they have to have their entire essence tracked by every company on the internet for that.

Regarding Nike and foot sizes specifically, rather than your general point: they do sell and ship shoes directly to consumers online, so yeah they’d have foot size info in that case, even if (as I hope) they don’t get that info from indirect purchases.
> IMO data should be radioactive for companies, especially if it approaches PII.

That's an idealistic, but highly unrealistic, thought.

As long as a market exists that can profit from exploiting PII, and is so large that it can support other industries, data will never be radioactive. The only way to make it so is with regulation, either to force companies to adopt fair business models, or by _heavily_ regulating the source of the problem—the advertising industry. Since the advertising industry has its tentacles deeply embedded everywhere, regulating it is much more difficult than regulating companies that depend on it.

So this is a good step by the EU, and even though it's still too conservative IMO, I'm glad that there are governments that still want to protect their citizens from the insane overreach by Big Tech.

> As long as a market exists that can profit from exploiting PII, and is so large that it can support other industries, data will never be radioactive.

The EU bureaucracy machine can be slow moving, but has the potential to fix this. The stricter the rules, the simpler the implementation. You could cut a LOT of the administrative burden by specifying what data is allowed to be stored at all, instead of what isn't.

Big tech needs to be put in their place, and as others have commented; if this kills your business model, your business model doesn't deserve to exist.

As a customer, I want the ability to choose the way in which I pay a business I interact with, with the consent of that business of course.

Europe gives me less control of my personal data than the US would. I am no longer allowed to decide that I'd rather choose services that take payment in data instead of services that take payment in Euros.

I think people who disagree with this perspective should be accommodated. It's a valid objection and technology inherently favors monopolies, so you can't really have the Facebook equivalent of a vegan restaurant or gay club. I'm not against forcing (large) tech companies to offer tracking-free plans at reasonable prices for those for whom this is the right tradeoff.

What Europe is doing is just plain stupid, though, and it will be felt most by those who can least afford it.

> I am no longer allowed to decide that I'd rather choose services that take payment in data instead of services that take payment in Euros.

Google, Microsoft and Apple don't really give you a choice, you will pay in Euros for your phone/PC, and then you will pay in your data as you use it whether you like it or not.

It should be perfectly fine if people want to pay with personal information, as long that personal information has zero social costs.

A prime example is sharing information about DNA since that has a social impact on relatives. Less obvious problem would be people in a position of social position, like say a judge or jury, since access to personal information in that situation provide unfair position of power in society. It also is a problem with voting, since access to voters personal information has a high risk of influence elections.

To take a more direct example, if you are paying your email provider with data, then you are also selling the information of anyone who send their emails to you. The sender is in an impossible position in that they can't know who the email provider is of a recipient (email forwarding is a thing), so the social cost is on the recipient if they sell the information.

> it will be felt most by those who can least afford it.

This sort of business model is problematic precisely because the poorest can't afford to refuse - that's a feature not a bug. Privacy is deemed a human right, and human rights shouldn't be for sale.

You could make the same argument supporting the legal sale of human organs, but as a society we've decided that kind of "payment" strips the poorest of their dignity and human rights.

The business model is inherently predatory for other reasons too. People see what they get right now - "free" access to the website they're on, but they're completely oblivious to the real costs because they're abstract, too many steps removed from each individual's actions, but they're very real and damaging in aggregate.

What you just described is actually only possible with data protection law. In Germany there are websites that ask you to accept ads and cookies or else u pay the monthly subscription fee. Without data protection law you likely just don't get the choice. Btw, the ads fee are priced in when you buy stuff.
I don’t understand where this is coming from. Isn’t Meta offering to EU users exactly the choice you are describing? (Even though in the case of the subscription we can’t really be sure they also don’t still use your data.)
The problem is that is also against the rules. GDPR bans using private data as a form of payment. You can give your data away freely but it can't be requested as a form of payment. In this case Meta is asking you to either pay with your money or your data. One is OK, the other isn't.
> data should be radioactive for companies, especially if it approaches PII

Cute theory. Fails in practice. Especially with LLMs on the horizon, this would be tantamount to unilateral nuclear disarmament. (Practically, it fails in that we haven't quantified the cost of breaches commensurate with what those of us who are security minded estimate it to be.)

I have advocated for privacy issues for a short while. "Data is radioactive" is the "defund the police" of our movement.

What do you mean? Why do LLMs need user data to operate?

On a side note, I also don't understand your comparison to "defund the police" - were there any places that fully applied it and demonstrated that it "fails in practice".

> Why do LLMs need user data to operate?

Training data?

> I also don't understand your comparison to "defund the police" - were there any places that fully applied it and demonstrated that it "fails in practice"

It's a famous example where a minority overreacting in a presentable way set the entire movement back.

Don't agree entirely, at least in italy you can tell chatgpt to not use your chats to train other models. Well, it's still going to be using memory if you tell it something, but whatever info you give theoretically should remain there and not be used afterwards
You made a claim but didn't backed it with an argument. How is banning corporate hoarding of user data similar to nuclear disarmament?
> How is banning corporate hoarding of user data similar to nuclear disarmament?

Nukes have clear downsides, ones one doesn't need a protractor or regression to prove. Our estimates of the costs of data breaches remain statistical.

> Especially with LLMs on the horizon, this would be tantamount to unilateral nuclear disarmament.

We should also be hoping for unilateral nuclear disarmament (I get your point on the infeasability though), but I don't see the parallels here. LLMs don't need personal data to work (I'd even imagine such data to be better off left out of the training data anyways, caveat for celebrities), and regardless of everything else whether the AI hypesters are to be believed about how world-changing AI/LLMs will be remains to be seen.

Also, as the OP article suggests, we can and are doing something about it. Things aren't perfect yet, but GDPR itself has already made huge waves and have made things better. From how I interpret this ruling, the dark pattern cookie banners are being scrutinized and are being put under the knife, so there's some hope that things will soon improve on that front.

> I have advocated for privacy issues for a short while. "Data is radioactive" is the "defund the police" of our movement.

Except we can already see a shift in the masses and their opinions here. People are becoming cognizant of the sheer amount of data all these tech companies harvest on them. I am consistently getting more and more of my non-technical-in-any-capacity friends asking me how to safeguard their data better, so I'm quite hopeful we're going to get there. All we need is to actually fucking hurt the FAANGs and their ilk. Cut the head off the snake and all that, if we actually hurt Meta as we should've a million times by now, then all the smaller players will automatically fall in line for fear of a similar world of hurt.

I don't see any parallels with unilateral nuclear disarmament and making exploiting user data unviable

>we haven't quantified the cost of breaches commensurate with what those of us who are security minded estimate it to be

We don't estimate GDPR violations as the true materialized damages either, we put a heavy % of yearly income per offense, large enough to deter it.

> We don't estimate GDPR violations as the true materialized damages either, we put a heavy % of yearly income per offense, large enough to deter it

Not remotely analogous to turning data into a liability. Particularly when the EU laws seem almost explicitly written to allow for offloading such risks to America and China.