So in your opinion, they are training on your data even if you toggle the "don't train on my data" checkbox off?
That's a bold assertion.
So in your opinion, they are training on your data even if you toggle the "don't train on my data" checkbox off?
That's a bold assertion.
Think of it as the Big Data hype some years ago.
Kek
More problematically there are camouflaged sharp spines pointed primarily in the direction of poorer people, and people not advised by lawyers.
But none of that matters here when the damaged parties include the megacorps of the world.
What they have been doing, with some narrow exceptions where they have lost billions of dollars in court cases*, is not at all obviously prohibited by copyright law. Neither web scraping (i.e. asking for copies of data from people you have every reason to believe are authorized to give you copies) or running algorithms on copyrighted data are generally copyright infringment. I say generally because the "algorithm" of "ctrl-c ctrl-v" is obviously an exception, and there's some argument that training is similar enough to be illegal - a fairly weak argument that is mostly losing in court but has some tiny chance of still succeeding.
The law doesn't have teeth to prohibit things not prohibited under the law - no matter how much many people would like them to be prohibited. This shouldn't be surprising.
Unlike with copyright, the law does pretty clearly prohibit violating contractual terms to not hang onto or use other peoples data for purposes other than the narrow ones laid out in the contract when you agreed to the contract.
* Namely acquiring copies of data from people who they know aren't authorized to make copies - i.e. torrenting.
So they are in fact literally putting copyrighted data into the model weights and reselling it.
Past behaviour informs future trust and I wouldn't trust these companies whatsoever.
Anthropic paid $1.5 billion for that, and never publicly deployed a model derived from the illegally downloaded data.
I'm not sure about the other companies off the top of my head - but I rather imagine they either never did this (I note that Google for instance already has lawfully acquired copies of basically every scrap of data you can imagine wanting to pirate) or are in the process of being sued or settled and I missed the news.
And none of this changes the fact that they did it in the first place and were comfortable doing so, thereby demonstrating that they are not trustworthy actors. If they could spend another 1.5B to advance their models with ill-gotten training data, there's every reason to believe they'd do it all over again.