back

by Panoramix·3y ago·view on hn ↗
It's not plagiarism at all. The AI is trained on 5 billion images yet it stores only 4gb of data. Thus it is impossible that it stores the actual work. For any image that the AI generates, you can't point to any image in the training data that the image is derived from.
4 comments
Why are you talking about AI?

This was about the humans consuming other people's content.

> Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessarily reimbursing the original source(s) of their creativity.

If humans make stuff that is too close to someone else's source materials then it is considered plagiarism and not "inspired by".

> For any image that the AI generates, you can't point to any image in the training data that the image is derived from.

Why can't you point to the Getty Images watermark that it is quite happy to reproduce? Isn't that surely evidence that it doesn't actually understand what it is reproducing?

> The AI is trained on 5 billion images yet it stores only 4gb of data. Thus it is impossible that it stores the actual work.

I have also seen billions of images, therefore I cannot be actually store the real images in my head and thus nothing I paint could ever be considered plagiarism. That's brilliant, I think there are a few law firms defending artists who would be looking to hire you.

How did they train the AI without first storing the data? It's not in the model, but it was used without permission in the pipeline that lead to creating that AI model.

I don't know if that counts as plagiarism, but there's clearly some use of this copyright material that the authors probably didn't envision and did not grant permission for. I have no idea what the law would be in cases like this

> How did they train the AI without first storing the data?

the data was originally permitted to be copied.

The question isn't whether the training is violating copyright - as long as the data set had permission to be viewed (which it must have, since it was public).

The question is whether the final result - the model/weights - is a derivative work of the training data set. If it is a derivative work, then the model must be in violation of copyright. But copyright law allows for sufficiently transformative work to be considered new, rather than derivative. So is training a model using methods like this constitute a transformative work?

Who's to say that GPT isn't just analogous, legally, to lossy compression algorithms, only with a particularly interesting interpretation of "lossy"?
There are some overfittings where you can point to source images, but that number is a lot smaller than 5 billion.