back
92 comments
I swear this kind of stuff is only going to get worse as people come to use and rely on ChatGPT and its ilk more broadly.

Dystopia? Idiocracy? I don't know, but I don't like it.

"This stuff" has been around for ages. Forty years ago in my teens, I read a book with a very memorable anecdote about locals insisting canabalism was "a local custom" and defending their right to practice it. The British official in charge replied "It's our custom to shoot cannibals."

When I was homeless, I asked around on the internet for a source for who said that. There is a real incident where a British official said something like "We shoot people who do that" but it wasn't about cannibalism.

But the toxic classist forum where I asked initially replied to me with basically "You're just a stupid homeless person misremembering that." No, I read it in my teens when I had a near photographic memory and was one of the top students in my high school class and everyone respected me as one of the smart people, long before the world decided I was some loser making things up. I'm quite clear the anecdote in the book was about cannibalism.

There is a real historical incident similar to it, but the book got the details wrong.

I have also read some crazy accounts of how Einstein's Theory was proved because the solar eclipse bent light so much, you could see a star that our Sun should have been obscuring.

Humans tend to believe things we read. If it's in writing, it has some kind of authority in our minds.

This is often not the case and we need to get better about recognizing that a lot of "writing" on the internet is just modern chit chat and not reliable.

> I have also read some crazy accounts of how Einstein's Theory was proved

FWIW, this was the observations during the solar eclipse of 1919, by Eddington and a bunch of less famous people. It apparently made headlines at the time.

My understanding is that the eclipse moved the apparent position of the star compared to where it would normally be observed, but not by very much.

All I'm saying is some accounts seem to really exaggerate how much this effect was.

It wasn't about bending of light though. The orbit of Mercury is slightly different from what Newton's theory would predict, and this was first observed during the eclipse.

The British officer story happened in Korea IIRC, the custom at the time wasn't cannibalism, but that women who lost their husbands would be killed to join them in the afterlife.

How did the people on the board know you were homeless and why do you keep randomly keep mentioning it?
It'll probably be a little bit like life before the web. You'd hear something and have no immediate way to verify the veracity of the claim. It's one of the reasons teachers pushed students to go to the library and use encyclopedias for citations when writing research papers.

Our species has obviously managed to make it pretty far without facts for the longest time. But we've comfortably lived with easily verified facts for 20-30 years and are now faced with a return to uncertainty.

If I had to guess, we'll see stricter controls on institutions such as Wikipedia that rely on credentialism and frequent auditing as a means to counter the new at-volume information creation capacity. But I don't really have the faintest idea of how this will turn out yet. It's wild to think about how much things are changing.

My teachers were always clear that encyclopedias were not to be used as a primary source either. They're not even a secondary source. Encyclopedias are tertiary sources.

They're better than Wikipedia... but only barely.

In the end you use Wikipedia and an encyclopedia the same way: to get a broad understanding of a topic as a mental framework, then look at the article's citations as a starting point to find actual, citable primary sources. (Plus the rest of the library's catalog/databases.)

exactly. information literacy starts with evaluating the sources. I have had numerous chats over the last few years where it's evident that people do not do due dillagence in their information gathering. it seems that either people aren't being taught this anymore or that they have given in to sloppy thinking.
Imagine a world where children have grown up, relying on ChatGPT for each and every question.
A world where children ask questions to unreliable entities who guess when they don't know the answer?

Pretty sure we just called it the 90s.

I would have killed to have ChatGPT growing up. It's amazing to have a patient teacher answer any question you can think of. GPT-4 is already far better than the answers you'll get on Quora or Reddit, and it's instant. So it's wrong sometimes. My teachers and parents were wrong plenty of times, too.
Apropos this, was tempted to submit https://www.youtube.com/watch?v=KfWVdXyPvWQ [1] after watching it last night, but maybe it's better to just leave it here, instead..

1: How A.I Will Self Destruct The Human Race (Camera Conspiracies channel)

Imagine that in five years from now, ChatGPT or one of its competitors will reach 98% factual accuracy in responses. Would you not like to rely on it for answering your questions?
chatcpt outputs everything just so confident since it's basically just a bullshit generator. it's markov chain word bots on steroids.
It’ll be a world where it’s important to know the right question to ask
Another problem to be solved with more computation and human backed reinforcement learning, surely.
How? As can be seen from these Citogenesis Incidents, humans cannot even tell when other humans are making up stuff that sounds like it could be real. How will ChatGPT, et al do it?
I just asked ChatGPT who made the first cardboard box, and it too believes the first story on this list: "The first cardboard box was invented by Sir Malcolm Thornhill in England in 1817. He created a machine that could make sheets of paper and then fold them into boxes."
No, it doesn't. It doesn't believe anything, it's just generating a story for you that sounds credible enough for you to go "yes, this is what an answer would look like". That's its job. That's its only job. Literally everything it says is fabricated, and if it happens to be the truth, that's a coincidence.
It effectively "believes" some things, as it will consistently emit certain statements in response to certain types of queries. It considers that information part of a good response.

There is information stored in its model. That information might not be correct.

Literally everything it says is fabricated

In what sense do you use the word "fabricated"? In the sense that it invented a falsehood with an intent to deceive, or in that it says things based upon prior exposure?

In 2016 I discovered that a mountain on the planet Ceres had officially been named "Ysolo Mons" after the Albanian festival that marks the start of the annual eggplant harvest.

There is no such festival. An anonymous Wikipedia editor had made it up and inserted it into Wikipedia's list of harvest festivals in 2012. Someone at NASA used the Wikipedia list for naming features on Ceres. (Ceres was the Roman goddess of agriculture.)

I wrote to the US geological survey to point this out. They changed the name of the mountain.

Full story on my blog: https://blog.plover.com/wikipedia/ysolo.html

Somehow reminds me of large language models. If they will be trained on data after the release of, say, GPT-3, they'll probably be trained on outputs of that model.
Yes, first thing that comes to mind: How much worse will this become with LLMs in the loop - almost assuming that this page was even submitted for that thought?
It's a major area of focus to improve the hallucination (or whatever the technically correct term is) of the model. I would bet we're pretty close to GPT actually evaluating sources for information and making judgements in how to weight those sources. I suspect this is going to upset a lot of people, and especially those in power.
Somehow I don't think this is going to be a problem. I can't exactly articulate why, but I'm going to try.

The success of an LLM is quite subjective. We have metrics that try to quantitatively measure the performance of an LLM, but the "real" test are the users that the LLM does work for. Those users are ultimately human, even if there are layers and layers of LLMs collaborating under a human interface.

I think what ultimately matters is that the output is considered high quality by the end user. I don't think that it actually matters if an input is AI generated or human generated when training a model, as long as the LLM continues producing high quality results. I think implicit in your argument is that the _quality_ of the _training set_ is going to deteriorate due to LLM generated content. But:

1) I don't know how much quality of the input actually impacts the outcome. Almost certainly an entire corpus of noise isn't going to generate signal when passed through an LLM, but what an acceptable signal/noise ratio is seems to be an unanswered question.

2) AI generated content doesn't necessarily mean it is low quality content. In fact, if we find a high quality training set yields substantially better AI, I'd rather have a training set of 100% AI generated content that is human reviewed to be high quality vs. one that is 100% human generated content but unfiltered for quality.

I don't necessarily think this feedback loop, of LLM outputs feeding LLM inputs, is necessarily the problem people say it is. But might be wrong!

LLM output cannot be higher quality than the input (prompt + training data). The best possible outcome for an LLM is that the output is a correct continuation of the prompt. The output will usually be a less-than-perfect continuation.

With small models, at least, you can watch LLM output degrade in real time as more text is generated, because the ratio of prompt to output in the context gets smaller with each new token. So the LLM is trying to imitate itself, more than it is trying to imitate the prompt. Bigger models can't fix this problem, they can just slow down the rate of degradation.

It's bad enough when the model is stuck trying to imitate its output in the current context, but it'll be much worse if it's actually fed back in as training data. In that scenario, the bad data poisons all future output from the model, not just the current context.

The 85% fatality rate for the water speed record had me go down a rabbit hole. The record hasn't been broken since 1978 (~315mph) and someone on reddit said it was because they stopped tracking the record due to so many deaths. I can't find any information online to corroborate this though.
Here's a fictitious citation that commonly appears on HN - "Dunning-Kruger effect":

> The expression "Dunning–Kruger effect" was created on Wikipedia in May 2006, in this edit.[1] The article had been created in July 2005 as Dunning-Kruger Syndrome. Neither of these terms appeared at that time in scientific literature; the "syndrome" name was created to summarise the findings of one 1999 paper by David Dunning and Justin Kruger. The change to "effect" was not prompted by any sources, but by a concern that "syndrome" would falsely imply a medical condition. By the time the article name was criticised as original research in 2008, Google Scholar was showing a number of academic sources describing the Dunning–Kruger effect using explanations similar to the Wikipedia article.[2]

[1] https://en.wikipedia.org/w/index.php?diff=55273744&diffmode=...

[2] https://en.wikipedia.org/wiki/Wikipedia:List_of_citogenesis_...

The spread of the "ranged weapon"/"melee weapon" classification terminology from the roleplaying games world into writings on real-world anthropology and military history (without acknowledgement of the direction of the borrowing!) is a personal pet peeve. I haven't been able to pinpoint Wikipedia, much less a specific article, as the source of this but it seems to have at the very least accelerated the trend.
Is anyone scholarly using "melee" that way? Or is it just ignorant amateurs? I've only encountered the latter (but I, and all my friends, find it hard to avoid saying "melee" to mean "hand-to-hand", because we were all D&D players before we were anything else).

Anyway, much as I do it, it annoys me too.

Related peeve, though as far as I know this is still restricted to gamers... How do you feel about "akimbo" meaning "wielding two guns, one in each hand", I believe that's from CounterStrike.

Or perhaps the word "glaive", to mean a thrown multi-bladed spinning weapon? I believe from Warcraft.

Does that mean the Dunning-Krueger effect has no basis in truth?
I'm from the internet and I can assure you it has no basis in truth.
ETA: see response

Probably not no basis, Dunning and Krueger really did so research & found [retracted] a negative correlation between self-rated ability and performance on an aptitude test afaik [/retracted]. But it's often overgeneralized or taken to be some kind of law rather than an observation.

A friend of mine intercepted one that was actually shared in other places: https://en.wikipedia.org/w/index.php?title=Said_the_actress_...
Reminds me of https://www.reddit.com/r/Jokes/comments/2wpf2h/indian_joke_a...

(sed /indian/native american/g)

> arbitrary addition to Coati, "also known as....the Brazilian aardvark"

But via citogenesis, the coati really became also known as the Brazilian aardvark. So the original claim is true, and this wasn't really citogenesis after all. More like self fulfilling prophecy.

In linguistics a sentence that causes itself to be true is called performative. (The standard example being "we are at war", spoken by someone with the authority to declare war.)

That doesn't really fit here, but it's a similar idea.

Deadspin saw one of these and made with a step-by-step summary: https://deadspin.com/how-espn-manufactures-a-story-colin-kae...
Many years ago I discovered and wrote the bullet in this article about the Zimbabwean hyperinflation:

"It is well known that Zimbabwe experienced severe hyperinflation in 2008..."

Very interesting article as a whole, of course.

The references in the comments suggest ChatGPT as providing this effect. But that is (or should be) unlikely, the "training" or moderation (tweaking?) should actually solve this problem. It should be relatively easy to separate it's own generation from sources. BUT where it will happen is when multiple instances of these language models compete with each other. ChatGPT quoting Bing or Bard output probably can't be reliably countered with internal training of ChatGTP, and the same goes for Bing & Bard and all the other myriad manifestations of these data mining techniques. (Unless they merge them togther?)
> BUT where it will happen is when multiple instances of these language models compete with each other.

That's what everyone else is saying already. Not sure what exactly you are arguing against.

Sorry a bit late replying-been away. It is not the competition that is bad, it that anything produced by them cannot be tweaked and so become "circular" sources. There will be no way to test for "truth", at least on an individual bot the training data can be tweaked to not use its own production as source. The competition will make the validity or "truth" of most data questionable. I guess it should be possible for an individual LMM to be trained for "truth" (reality?) but it becomes almost impossible for a LMM to discern truth when the sources it is analyzing are of generated by another LMM
> Terms that became real

I suppose that was inevitable.

Really hoped that the fact this came from xkcd was itself an example of citogenesis.