back

by CrypticShift·3y ago·view on hn ↗
Without all that human energy and enthusiasm on the web over a generation, this level of "AI intelligence" would not have been possible in the first place.

So, here is the paradox: (in the future) the less we will continue to create content (HN, Blog...), the less ChatGPT-like answers will be "relevant". So, we may go back to asking our fellow humans more.

What I'm more worried about is all that cheap (pricewise and quality-wise) AI content proliferation, because it will worsen the effects of that Human-content decrease. It will be AI feeding on itself, like zombies.

1 comments
Why can't AI learn from other AI? Won't be long before they're upvoting each other and telling us to go away.

"Such a low quality post. Are you human or something? Go back to reddit!"

This is only a very rough analogy, but consider these text AIs like lossy JPEG compression for images -- using an efficient statistical representation of the content, you can store a lot of information in a little space and reconstruct something aaaalmost like the original. If I (and a few hundred million other people) have gone through the trouble to create a nice PNG picture or whatever, and then someone comes along and saves it as a JPEG at.. say.. 95% quality, that JPEG will look pretty great. Hard to distinguish from its source.

But what happens when our JPEG saver only has other JPEGs to start with? And soon, is saving copies of copies of... you get the idea. The images lose more and more of their original real-world structure, and become dominated instead by artifacts of the compression and reconstruction processes.

That's basically the (a) challenge for this AI training paradigm of scraping web-scale content. The more that content is generated by the previous generations of AI, the more we're asking the next generation to learn from copies of copies of the original human intelligence that powers them

That assumes AI can only make ever-degrading copies of copies and are completely unable to synthesize new information out of existing information. I don't think is true any more (even just at today's capabilities).

It is producing novel information -- pictures or text or otherwise.

Besides, human creativity isn't borne of a vacuum either. Most humans works are derivatives of some other work and depend on the creator's life experiences, etc. Our entire education system is based on "people teaching other people what they know", at a lower fidelity, copies-of-copies style.

Like I said, rough analogy. The "lossy compression" here is in the implicit world models these networks learn, not necessarily in the generation space (though we certainly see deviations from reality there too). However, the bigger difference between learning from SEO-generated blogspam and human teaching is the lack of feedback mechanisms to correct errors in generated output that then get taken as ground truth.
> assumes AI ... are completely unable to synthesize new information out of existing information

I love it when we go back to assumptions. This is exactly what I'm assuming. However, the problem is not solved. what exactly is "new" information? here is more assumptions for you right there.

> Most humans works are derivatives of some other work

Good point. "Most", however, does not equal "All". collectively, we are not just copying. there is a newness. (again, what is new?)