back

by toddmorey·1y ago·view on hn ↗
No way OpenAI will ever “good citizen” this. Tools to opt out of training sets will only come if they are legally compelled. Governments will have to make respecting some sort of training preference header on public content mandatory I think.

The fact that photographers have to independently submit each piece of work they wanted excluded along with detailed descriptions just shows how much they DONT want anyone excluding content from their training data.

5 comments
Reminds me of the time when p2p music sharing became popular and the record companies had to submit every song they did not want to get shared along with an explanation to every person who had Napster installed.

Or was it that the record companies got to sue individuals for astronomic amounts of made up damages for every song potentially shared?

Which one was it?

> and the record companies had to submit every song they did not want to get shared along with an explanation to every person who had Napster installed.

...when did that ever happen? The post-Napster-but-pre-BitTorrent era (coincidentally, the same time-period as the BonziBuddy-era) was when Morpheus, KaZaA, eDonkey, Limewire, et cetera were relevant, and they got-away-with-it, in-part, by denying they had any ability to moderate their users' file-sharing; there was no "submitting of every song" to an exclusion-list because there was no exclusion-list or filtering in the first place.

This almost got me too, but you missed GP’s point. See their next paragraph beginning “Or”. (It's about double standards for individual versus corporate copyright infringement.)
100%, like most opt-outs this exists as a checklist feature that proponents can point to and hopefully convince bystanders. You muddy the waters by allowing someone to with great effort technically possibly achieve the thing they want, maybe, for now, until you close it in 2 years and everyone says "well that makes sense nobody used that feature anyways".
> The fact that photographers have to independently submit each piece of work they wanted excluded along with detailed descriptions just shows how much they DONT want anyone excluding content from their training data.

That's bloody brilliant. If you don't want us to scrape your content, please send us your content with all of the training data already provided so we will know not to scrape it if we come across it in the wild. FFS

The tech industry’s understanding of consent is terrifying.
Understanding is a curious choice of words. I’d have gone with total disregard
Or even contempt for?
Mirrors that of a sexual predator.

"Oh I'm not groping you today? No worries, I'll be back tomorrow."

> "Oh I'm not groping you today? No worries, I'll be back tomorrow."

the trick is to come back tomorrow, but with a rusty and jagged metal mousetrap hidden in one's underwear... and a camera for posterity, and some witnesses to come point-and-laugh at the perp.

This is something I frequently point out to. If someone understood constent like tech companies do, they'd be banned from a few bars. Look up "rules for consent" and think about how consensual your relationship with tech companies is.

Here's one: https://stopsexualviolence.iu.edu/policies-terms/consent.htm...

It mirrors the rest of society's lack of understanding of consent. Sunrise, sunset
They learned from Google, who to this day requires you to suffix your wifi network name with _NOMAP if you do not want it to be used by their mapping services.
Sounds like they want photographers to do the data labeling for them....
Insofar as data for diffusion / image / video models are concerned, the rise of synthetic data and data efficiency will mean that none of this really matters anyway. We were just in the bootstrapping phase.

You can bolt on new functional modules and train them with very limited data you acquire from Unreal Engine or in the field.

I don’t entirely agree. For example, it’s a very popular scheme on Etsy right now to use LLMs to generate posters in the style of popular artists. Any artist should be able to say hey I don’t want my works to be part of your training set to power derivative generations.

And I think it should even apply retroactively so that they have to retrain their models that are already generating works from training data consumed without permission. Of course, OpenAI would fight that tooth & nail but they put themselves in this position with a clear “take first ask permission later” mentality.

Dumb question: Why does Etsy allowed clearly reproduced/copied works? AI or not.

Like selling it for money seems like a clear line crossed, and Etsy is the perfect gatekeeper here.

Etsy stopped caring a while ago, it was supposed to be a marketplace specifically for selling handmade items but they allowed it to be overrun with mass produced tat dropshipped direct from the factory. Turning a blind eye to plagiarism with or without AI is just the next logical step from there.
> Why does Etsy allowed clearly reproduced/copied works

They don't, in that they'll ban you for it once you're big enough

Style isnt protected?
Impossible to put the toothpaste back in the tube.
A motivated legislature with skilled enforcement personnel could get the toothpaste back in the tube in short order provided they displayed an anomalous insensitivity to making the money sad.
When has something like this ever happened? It feels like legislature exists to make money happy.
One big example involves making a lot of US plantation-money sad and scared, so much that it started a civil-war.

Granted, that was more the exception than the rule...

> When has something like this ever happened?

Anything that used to be freely available but no longer is. Once upon a time Laudanum (tincture of opium) used to be the OTC painkiller of choice. In slightly more recent times, there's asbestos. In certain locales, gambling. There's countries that have reigned in lootboxes.

> It feels like legislature exists to make money happy.

Come on now, it doesn't just "feel" that way, you know for a fact that is indeed the purpose of the modern US legislature.

Should any artist be able to tell another artist: hey don't copy my work when you're learning, I don't want competition?

It seems like they are deeply upset someone has figured out a way for a machine to do what artists have been doing since time immemorial.

There are two major differences between art generators and human artists:

1) human artists are legal persons and capable of being held liable in civil court for copyright infringement; having a machine with no legal standing do the copyright infringement should be forbidden because it is difficult to detect, impossible to avoid, and a legal nightmare to unravel.

2) human artists are capable of understanding what flowers, Jesus on the cross, waterfalls, etc actually are, whereas DALL-E is much dumber than a lizard and not capable of understanding these things, so using the verb "learning" to describe both is extremely misleading. DALL-E is a statistical process which is barely more sophisticated than linear regression compared to a human brain. It is plain wrong to say stuff like this:

> It seems like they are deeply upset someone has figured out a way for a machine to do what artists have been doing since time immemorial.

when nobody has even come close to figuring that out! If DALL-E worked like a human artist it would know what a bicycle is: https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_pr... But it doesn't. It is a plagiarism machine that knows how to match "bicycle" with millions images having a "bicycle" tag, and uses statistics to smooth things together.

Llms are not humans and shouldn’t be anthropomorphized as a strategy to get around copyright infringement.
Absolutely correct, LLMs are code. If that code ingests data without precisely replicating it, that's fair use and the end of the discussion.
And if they do get anthropomorphized... then the people in charge of that company need to be charged with the heinous crime of enslaving children.
I've been uhh.. suffering(?) a different perspective. A company may hire a human to talk for it or its owner may talk for it.. the server rack and the software are things it can own. I don't think others are nesasarly any longer talking for it.

It is an absurd leap we've made but companies are also legal persons.

The companies are still of human design full of human behaviour and human characteristics while the LLMs actively try to imitate humans.

The dictionary saying: anthropomorphized: attribute human characteristics or behaviour to (a god, animal, or object).

If it passes the Turing test surely anthropomorphizing is fair game?

(I have no stake in this)

That's an opinion that will be tested in a court soon enough.
My mom said the teachers in her painting classes would have the students recreate works and were very clear on which artists had given permission for those derivative works to be sold. Others they could only admire at home.

The problem is not "when learning", the problem is "when distributing". Courts will determine whether or not disseminating or giving access to a model trained on protected works counts as distributing protected derivative works or not.

> The problem is not "when learning", the problem is "when distributing".

Technically making a copy to bring home for your own use is also problematic, just much less likely to get you into trouble. (Still a step removed from learning the skills and technique of making a copy, however.)

This analogy seems to be made every time this comes up on HN, but I don't think it really holds water. First of all, when a human artist learns from another, it's inherently a level playing field for competition; the junior and senior are both human, neither are going to be 1,000,000x more productive as the other. So the senior artist really doesn't have that much to worry about. And the senior artist recognizes that they themselves were once a junior and had to learn from their seniors, so it's a debt paid forward which results in more art. And becoming a master artist or even an imitator takes dedication and hard work, even with lots of artwork to learn from, so that keeps competition to a certain level.

When it takes decades to develop an art style that a machine can copy in days, and then churn out derivative variations in seconds, it's no longer a level playing field. The machine can dramatically under-cut the artist who developed their style, much more than a copycat human artist could. This does become not just a threat to the livelihoods of artists, but also a disincentive to the development of new art styles.

In this case, patent law may be an apt comparison for the world we're entering. Patent law was developed with the idea in mind that it is a problem if a human competitor could simply take an invention, learn how it works, and then mass produce copies of it. There are several reasons for this, including creating an incentive for technology development, and also expediently transitioning IP to the public domain. But patents were added to the legal system basically because otherwise an inventor would not be on a level playing field with the competition, because it takes so many more resources to develop a new invention than to produce clones.

Existing IP law was built in a world where it was believed that machines were inherently incapable of learning and mass-producing new artistic works using styles learned from artists. It was not necessary to protect artists from junior artists learning how to work in their style, as long as it wasn't a forgery. But in a world of machine learning, perhaps we will decide it's reasonable to protect artists from machine copycats, just like we decided it was reasonable protect technology inventors from human copycats.

The patent system is not the right implementation; it's expensive to file a patent, and you need skilled lawyers to determine novelty, infringement, and so on. But for art and machine learning, it might be much simpler: a mandatory compensation for artists' work used as training data. Something like this is sometimes used in the music industry to determine royalties for radio broadcasting, or to account for copies spread by file sharing.

Surely all this applies to code and the written word in general?

People allowed (and encouraged) read access to websites so Google would index and link. Now Google et al summarise and even generate. All of that is built on our collective output. Surely everyone deserves a cut? The free sharing licenses that were added to repos didn’t account for LLM’s, so we should revisit it so all creators get their dues, not just those who traditionally got paid.

(Yes, for what it's worth I agree with this!)
When a human does it, it's fine.

When OpenAI's servers do it its copyright infringement.

We don't apply copyrights to human brains, but we do apply copyright to computer memory.

not sure if to pull out my racism card or to help burn the printing press.
I dont know why you are getting downvotes, you are absolutely correct.

A training method for some authors who want to adopt an older artists voice is to literally rewrite their novels. Word for word. They will go through an entire authors catalogue and reproduce them, so that they can learn to mimic them when creating something new.

You go ahead and automate the process, and suddenly the world is ending.

Ditto all other kinds of art. Heck I knew of 3 living artists doing this to each other in real time.

HN is not only filled with people who never learned how to code, it's filled with people who never learned how to write either.

Hunter Thompson literally sat down and typed out every word of Hemingway's novels so he could figure out what good writing feels like.

Why is he allowed to do it in private, but an LLM isn't?

Has synthetic data become a big part of image/video models?

I understand why it's useful and popular for training LLMs, but I didn't think it was applicable to generative image/video work.

I haven't had the chance to train diffusion models but for detection models synthetic data is absolutely how you get state of the art performance now. You just need a relatively tiny extremely high quality dataset to bootstrap from.
For clarity, I do agree that synthetic data is huge for training AI to do certain tasks or skills. But I don’t think creative work generation is powered by synthetic data and may not be for a quite while.
Isn't that just weird cope? I mean, why not just LLM automate UE if that's the goal & how isn't that itself going to get torpedoed by Epic?