The fact that photographers have to independently submit each piece of work they wanted excluded along with detailed descriptions just shows how much they DONT want anyone excluding content from their training data.
Or was it that the record companies got to sue individuals for astronomic amounts of made up damages for every song potentially shared?
Which one was it?
...when did that ever happen? The post-Napster-but-pre-BitTorrent era (coincidentally, the same time-period as the BonziBuddy-era) was when Morpheus, KaZaA, eDonkey, Limewire, et cetera were relevant, and they got-away-with-it, in-part, by denying they had any ability to moderate their users' file-sharing; there was no "submitting of every song" to an exclusion-list because there was no exclusion-list or filtering in the first place.
That's bloody brilliant. If you don't want us to scrape your content, please send us your content with all of the training data already provided so we will know not to scrape it if we come across it in the wild. FFS
"Oh I'm not groping you today? No worries, I'll be back tomorrow."
the trick is to come back tomorrow, but with a rusty and jagged metal mousetrap hidden in one's underwear... and a camera for posterity, and some witnesses to come point-and-laugh at the perp.
Here's one: https://stopsexualviolence.iu.edu/policies-terms/consent.htm...
You can bolt on new functional modules and train them with very limited data you acquire from Unreal Engine or in the field.
And I think it should even apply retroactively so that they have to retrain their models that are already generating works from training data consumed without permission. Of course, OpenAI would fight that tooth & nail but they put themselves in this position with a clear “take first ask permission later” mentality.
Like selling it for money seems like a clear line crossed, and Etsy is the perfect gatekeeper here.
They don't, in that they'll ban you for it once you're big enough
Granted, that was more the exception than the rule...
Anything that used to be freely available but no longer is. Once upon a time Laudanum (tincture of opium) used to be the OTC painkiller of choice. In slightly more recent times, there's asbestos. In certain locales, gambling. There's countries that have reigned in lootboxes.
> It feels like legislature exists to make money happy.
Come on now, it doesn't just "feel" that way, you know for a fact that is indeed the purpose of the modern US legislature.
It seems like they are deeply upset someone has figured out a way for a machine to do what artists have been doing since time immemorial.
1) human artists are legal persons and capable of being held liable in civil court for copyright infringement; having a machine with no legal standing do the copyright infringement should be forbidden because it is difficult to detect, impossible to avoid, and a legal nightmare to unravel.
2) human artists are capable of understanding what flowers, Jesus on the cross, waterfalls, etc actually are, whereas DALL-E is much dumber than a lizard and not capable of understanding these things, so using the verb "learning" to describe both is extremely misleading. DALL-E is a statistical process which is barely more sophisticated than linear regression compared to a human brain. It is plain wrong to say stuff like this:
> It seems like they are deeply upset someone has figured out a way for a machine to do what artists have been doing since time immemorial.
when nobody has even come close to figuring that out! If DALL-E worked like a human artist it would know what a bicycle is: https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_pr... But it doesn't. It is a plagiarism machine that knows how to match "bicycle" with millions images having a "bicycle" tag, and uses statistics to smooth things together.
It is an absurd leap we've made but companies are also legal persons.
The companies are still of human design full of human behaviour and human characteristics while the LLMs actively try to imitate humans.
The dictionary saying: anthropomorphized: attribute human characteristics or behaviour to (a god, animal, or object).
If it passes the Turing test surely anthropomorphizing is fair game?
(I have no stake in this)
The problem is not "when learning", the problem is "when distributing". Courts will determine whether or not disseminating or giving access to a model trained on protected works counts as distributing protected derivative works or not.
Technically making a copy to bring home for your own use is also problematic, just much less likely to get you into trouble. (Still a step removed from learning the skills and technique of making a copy, however.)
When it takes decades to develop an art style that a machine can copy in days, and then churn out derivative variations in seconds, it's no longer a level playing field. The machine can dramatically under-cut the artist who developed their style, much more than a copycat human artist could. This does become not just a threat to the livelihoods of artists, but also a disincentive to the development of new art styles.
In this case, patent law may be an apt comparison for the world we're entering. Patent law was developed with the idea in mind that it is a problem if a human competitor could simply take an invention, learn how it works, and then mass produce copies of it. There are several reasons for this, including creating an incentive for technology development, and also expediently transitioning IP to the public domain. But patents were added to the legal system basically because otherwise an inventor would not be on a level playing field with the competition, because it takes so many more resources to develop a new invention than to produce clones.
Existing IP law was built in a world where it was believed that machines were inherently incapable of learning and mass-producing new artistic works using styles learned from artists. It was not necessary to protect artists from junior artists learning how to work in their style, as long as it wasn't a forgery. But in a world of machine learning, perhaps we will decide it's reasonable to protect artists from machine copycats, just like we decided it was reasonable protect technology inventors from human copycats.
The patent system is not the right implementation; it's expensive to file a patent, and you need skilled lawyers to determine novelty, infringement, and so on. But for art and machine learning, it might be much simpler: a mandatory compensation for artists' work used as training data. Something like this is sometimes used in the music industry to determine royalties for radio broadcasting, or to account for copies spread by file sharing.
People allowed (and encouraged) read access to websites so Google would index and link. Now Google et al summarise and even generate. All of that is built on our collective output. Surely everyone deserves a cut? The free sharing licenses that were added to repos didn’t account for LLM’s, so we should revisit it so all creators get their dues, not just those who traditionally got paid.
When OpenAI's servers do it its copyright infringement.
We don't apply copyrights to human brains, but we do apply copyright to computer memory.
A training method for some authors who want to adopt an older artists voice is to literally rewrite their novels. Word for word. They will go through an entire authors catalogue and reproduce them, so that they can learn to mimic them when creating something new.
You go ahead and automate the process, and suddenly the world is ending.
Ditto all other kinds of art. Heck I knew of 3 living artists doing this to each other in real time.
Hunter Thompson literally sat down and typed out every word of Hemingway's novels so he could figure out what good writing feels like.
Why is he allowed to do it in private, but an LLM isn't?
I understand why it's useful and popular for training LLMs, but I didn't think it was applicable to generative image/video work.