back
233 comments
This reminds me of something from “The language instinct” by Steven pinker. One thing that stuck with me was a story of a girl named Denys who is or was severely mentally handicapped, but speaks perfectly.

She can’t read, write, handle money, but confabulates engaging and convincing stories about receiving mail, haggling with her bank, her boyfriend. She speaks with proper grammar, intonation, etc, but doesn’t demonstrate any understanding of what she talks about beyond the words. She doesn’t have a bank account or a boyfriend.

Trying to find more and I can’t find anything about her on the internet, I would have to look at the references. Maybe it’s questionable.

Anyway, I like to think about what the hell we even mean when we say we “know something”, and where we should actually put the bar for AI. Where is the line between confabulation and knowledge? Is knowledge just confabulation that happens to be correct?

I feel like it’s possible that Denys could feel that she “knows” what she is saying, but is simply trapped in, with a working knowledge of the world and is just unable to translate it into actions, and words are just her only degree of freedom.

I don’t remember where I was going with this, but I agree that “confabulation” is a much better term to use, but we also shouldn’t discount what it means to be able to confabulate.

Some of the most annoying moments in my career were occasional coworkers who talked a lot without making a point. They just invoked topics loosely related to the problem at hand, like stories from previous jobs or what they've read on the internet about this problem, until the manager said something like "are you trying to say that..." and then formulated what they thought the point was, except that the speaker clearly never intended to say anything like that. All the other engineers in the room are giving each other side eye but the manager thinks they've just got "a valuable input".

The tactic is simply to talk relevantly until the listener fools themselves into thinking they understand you, except that "understanding" will be an illusion and 100% manager's own analytical effort. And it's honestly embarrassing how often and well that works. Quite a few of those guys were promoted first, too.

I'm pretty sure our upcoming future with AI will be like that quote from forest ranger at Yosemite National Park, on why it is hard to design the perfect garbage bin to keep bears from breaking into it: "There is a considerable overlap between the intelligence of the smartest bears and the dumbest tourists."

The whole idea of a criteria to separate humans from AI is a misconception. I think any human who confuses smooth talking with making sense will be ruthlessly exploited in relatively near future. There are humans who already do this successfully to other humans, what do you think will happen when they get to automate their efforts.

I had the same thought that "hallucinate" is too strong a term for what LLMs are doing.

Instead I would say they DREAM. Dreams follow some logic but basically they are disjoint from reality. They involve same characters as our real-life experiences but what those characters do in a dream is not based on reality.

So one could think that ALL LLMs ever do is dream but much of the time their dreams are dreams which feel very real.

Our dreams are basically an LLM based on our real experiences but recombined in arbitrary ways, still retaining some logical structure but now detached from reality because when you recombine different experiences including our thoughts from real life, the result cannot really reflect reality very well. LLMs are better in this respect than our dreams, but sometimes their "dreams" get really obviously detached from reality.

Or if you want to use the term "hallucinate" with LLMs then fine but its more like they hallucinate all the time, it's just that sometimes, even often, their hallucinations seem to agree with the reality so well that we cannot tell the difference.

LLMs do not describe reality because they cannot experience reality, all they can give us is an "average description" created from many existing descriptions.

One way of looking at it is to say that all LLMs tell us is hearsay.

There's Lots of evidence indicating language models roughly know when they hallucinate.

GPT-4 logits calibration pre RLHF - [https://imgur.com/a/3gYel9r](https://imgur.com/a/3gYel9r)

Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - [https://arxiv.org/abs/2305.14975](https://arxiv.org/abs/2305...

Teaching Models to Express Their Uncertainty in Words - [https://arxiv.org/abs/2205.14334](https://arxiv.org/abs/2205...

Language Models (Mostly) Know What They Know - [https://arxiv.org/abs/2207.05221](https://arxiv.org/abs/2207...

It seems very much like hallucinations aren't a randomness or representation issue. The computation already knows. It just doesn't care about telling you this

I knew someone who was like Denys after a stroke. They spoke well, in grammatically correct sentences that made sense even if the content was not rooted in reality. There was not much coherence between sentences or in longer statements but it was more like quick gradual shifts not hard jumps from topic to topic.

So not unlike early LLM chat in areas where the model had not a lot of data about.

It'd be well worth verifying the story because if Steven Pinker made an error like that it's a fairly big deal.
Here's the full passage in The Language Instinct about Denyse:

    Here is another interview, this one between a fourteen-year-old girl called Denyse and the late psycholinguist Richard Cromer; the interview was transcribed and analyzed by Cromer's colleague Sigrid Lipka.

    > I like opening cards. I had a pile of post this morning and not one of them was a Christmas card. A bank statement I got this morning!

    < [A bank statement? I hope it was good news.]

    > No it wasn't good news.

    < [Sounds like mine.]

    > I hate . . . , My mum works over at the, over on the ward and she said "not another bank statement." I said "it's the second one in two days." And she said "Do you want me to go to the bank for you at lunchtime?" and I went "No, I'll go this time and explain it myself." I tell you what, my bank are awful. They've lost my bank book, you see, and I can't find it anywhere. I belong to the TSB Bank and I'm thinking of changing my bank 'cause they're so awful. They keep, they keep losing . . . [someone comes in to bring some tea] Oh, isn't that nice.

    < [Uhm. Very good.]

    > They've got the habit of doing that. They lose, they've lost my bank book twice, in a month, and I think I'll scream. My mum went yesterday to the bank for me. She said "They've lost your bank book again." I went "Can I scream?" and I went, she went "Yes, go on." So I hollered. But it is annoying when they do things like that. TSB, Trustees aren't. . . uh the best ones to be with actually. They're hopeless.

    I have seen Denyse on videotape, and she comes across as a loquacious, sophisticated conversationalist—all the more so, to American ears, because of her refined British accent. (My bank are awful, by the way, is grammatical in British, though not American, English.) It comes as a surprise to learn that the events she relates so earnestly are figments of her imagination. Denyse has no bank account, so she could not have received any statements in the mail, nor could her bank have lost her bankbook. Though she would talk about a joint bank account she shared with her boyfriend, she had no boyfriend, and obviously had only the most tenuous grasp of the concept "joint bank account" because she complained about the boyfriend taking money out of her side of the account. In other conversations Denyse would engage her listeners with lively tales about the wedding of her sister, her holiday in Scotland with a boy named Danny, and a happy airport reunion with a long-estranged father. But Denyse's sister is unmarried, Denyse has never been to Scotland, she does not know anyone named Danny, and her father has never been away for any length of time. In fact, Denyse is severely retarded. She never learned to read or write and cannot handle money or any of the other demands of everyday functioning.

    Denyse was born with spina bifida ("split spine"), a malformation of the vertebrae that leaves the spinal cord unprotected. Spina bifida often results in hydrocephalus, an increase in pressure in the cerebrospinal fluid filling the ventricles (large cavities) of the brain, distending the brain from within. For reasons no one understands, hydrocephalic children occasionally end up like Denyse, significantly retarded but with unimpaired—indeed, overdeveloped—language skills. (Perhaps the ballooning ventricles crush much of the brain tissue necessary for everyday intelligence but leave intact some other portions that can develop language circuitry.) The various technical terms for the condition include "cocktail party conversation," "chatterbox syndrome," and "blathering."
Source: http://f.javier.io/rep/books/The-Language-Instinct-How-the-M...
That story actually cracked me up, the way he ends it.
"Hallucinate" predates LLMs.

For instance, before transformers, when we finetuned ESRGAN for specific styles/functions, we used to say it would "hallucinate" in detail that couldn't possibly be in the original image pixels. GenAI wasn't much of a thing back then, and we never thought of the small GAN models as "remembering" things even though thats common language in transformer models. It felt more like we were teaching the GAN a job or style, not encoding memory.

And I think it fit! "Hallucinating" detail into a blurry pixel blob feels like a more accurate analogy to the human condition, and we weren't at the point where it was "confabulating" and generating big objects out of the blue like a diffusion model can.

Sure. Yet, there were "deep dreams" to generate "unreal" images back then. In my opinion, these are hallucinations - things that happen in dreams (or under the influence of various substances).

See https://blog.research.google/2015/06/inceptionism-going-deep... and https://distill.pub/2017/feature-visualization/.

I would differentiate these dreamy generations (i.e., things that we don't see ordinarily) from generations that are incorrect. Sure, there is overlap.

I'm inclined, perhaps mistakenly, to think of it as similar to how our brains fill in our peripheral vision, to respond to the lack of data.

https://neurosciencenews.com/peripheral-vision-brain-illusio...

Right, so basically "hallucinate" fits better for visual generative ML and "confabulate" for text. Seems like a sensible distinction to me.
karpathy says may have coined it in his RNN post https://twitter.com/DrJimFan/status/1703072983903060260
While we're talking about terminology, I think both "hallucinate" and "confabulate" are misleading and lead to inappropriate intuitions about how language models work.

Both words imply that there's something erroneous or anomalous happening when an LLM emits incorrect facts, as if the model can be in either a correct or incorrect state. But from a system perspective, there's no difference between the two scenarios: the model is consistently generating the most probable output based on its weights and context. It doesn't sometimes confabulate... it always confabulates.

The concept of "truth" or "falsehood" are something purely external to the model that humans bring when evaluating the output. They aren't part of the model itself.

Yea IMO it’s more useful to think of AI from a purely tech pov like this. The important thing is that brains resemble tech not that tech resembles a brain.
I like this a lot.

I recall seeing a post on here that roughly just said, "stop saying hallucinate", which sure, is technically correct ("the best kind of correct"), but not super helpful. Contrast that with this post, which offers a constructive suggestion: replace it with "confabulate".

I agree that C is an improvement over H, for different reasons from the author. Here's why: From a public comms standpoint, the big problem with saying "hallucinate" is that it leads to misconceptions, because people think they know what "hallucinate" means, and may anthropomorphize their pop psych grasp of H onto the LLM.

Just like H, "C" seems to have a technical meaning in the psych literature, but it's far less in common parlance than "H", so lay readers have no fixed idea of what exactly "C" means. Nobody goes around experimenting with "confabulagens" to seek a spiritual high. Also, "confabulate" retains a non-clinical meaning that's similarly obscure, from long before the psychologists adopted it, and it's pretty evocative of what's really going on here (e.g. via related words like "fable" and "fib"). So "confabulate" seems ripe for adoption to describe this specific computational phenomenon, as we've done to so many other mundane words in the past (deprecate, and so forth).

If the psychiatric "C" phenomenon is also a better analogy for LLM behavior, as OP claims (and some posters dispute?), so much the better. But my main reason for liking C, is that it's less likely than H to make people unduly confident that they understand it.

(But even if the community came to a consensus at this late stage that "C" is a better term, the subsequent problem of updating every stale paper and post that uses the old term would run us up against another doozy: cache invalidation.)

I've been using confabulate privately for a while about this phenomenon, because if you know anything about dementia (my grandmother had it - sitting in the car with her you got used to the "covers" feeling like chatbot answer dodges to some extent) or split brain experiments, then it's quite a remarkable phenomenon that does feel a lot more like what an LLM is doing.

Hallucinate generally implies experiencing sensory inputs which aren't there, but having a rational reaction to them - in the popular psyche though, people who hallucinate are still "crazy" - unfairly so, because you can function normally with hallucinations under quite a number of circumstances (i.e. I remember someone saying that they realized that if they saw people without faces, it was fine because they weren't really there - apparently quite common, since face recognition is a different part of the brain).

Basically: reacting to invalid sensory inputs implies rational behavior, but to non-real inputs.

Confabulation on the other hand is different because it's closer to "trying to rationalize an autonomous behavior without knowledge of the sensory inputs causing it" (maybe, it's a bizarre phenomenon). In split-brain patients in lab settings, it manifests as being unable to answer a question about why you're having a physiological response to imagery which is being shown only to the right-side of the brain, whereas language is processed on the left - but rather then be confused, people will apparently make something up that they're unaware is actually a "lie" (in quotes because, well, they're not lying - as far as we know there's no intent or even knowledge that is a lie).

Where this leaves us with LLMs I don't really know, because it feels like an imperfect descriptor: except perhaps with the context that LLMs are fairly limited networks, and so to some extent the whole "being wrong" phenonmenon is in fact just a failure of the attention mechanism to be able to draw the right data together.

Indeed it seems like replacing "hallucinate" with "confabulate" would improve understanding of the phenomenon in english, but it could cause a problem in Romance languages.

Indeed, the Latin verb "confabulor has the meaning of "discuss, converse" [0], not specifically "inventing things", and Romance languages already have words derived in sound and meaning from this classic Latin sense. For instance French already has "confabuler" meaning "speak familiarly with someone" [1]. Thus in Romance languages, this proposed usage of "confabulate" imported from the psychiatric word could be understood poorly due to preexisting similar-sounding words that do not include the notion of "making things up".

Just proposing thoughts here, not advocating for a final decision. Maybe it's still good to adopt this wording!

[0] : https://en.m.wiktionary.org/wiki/confabulor#Latin [1] : https://www.dictionnaire-academie.fr/article/A9C3479

I don't agree with the author's argument, because LLMs sometimes give wrong information even though they have access to the correct information. For example, querying an LLM about a fictional story will often elicit false answers. But if you respond with 'x is incorrect, now tell me what actually happened' you'll often get an apology and the correct answer. My informal sense is that this happens more with how/why questions than who/what/where/when ones.
LLMs do not confabulate or hallucinate. They predict the next word in a sequence based on correlations derived from a lot of training data. Sometimes that data happens to be structured enough to statistically generate a word that is fluent and corresponds to what we as readers understand as reality. Sometimes, due to the unclear nature of "truth" and the necessarily incomplete data used to train the model, the word or words that are statistically generated are merely fluent but do not correspond to reality.

The answer to "What are LLMs but humans with extreme amnesia and no central coherence?" is that they are not like humans at all, and only resemble them superficially due to our anthropomorphic tendency to infer a mind upon seeing fluent and roughly correspondent text.

I'm not sure if confabulate is exactly what LLMs do (though it seems closer than the implications of hallucinate).

But neurotypical, neurodiverse, perfectly functional people, etc. all confabulate or do something similar on a regular basis, in verbal and written mediums, and often do so in good faith. It's human instinct to communicate, even if you are uncertain and unaware of the full context of the discussion.

Teachers, customer service reps, executives, shop keepers, doctors, nurses, domain experts, authors of textbooks, it doesn't matter who it is, they'll probably confabulate or equivocate or do some other type of communication that isn't immediately useful. Yet it's still a useful activity to just talk to someone or read a less than rigorous book for the purposes of learning (discounting the relationship forming part, which is also useful). And so is using LLMs, even for casual users. So long as they understand that limitation, whether its with a chatbot or a real person. Not everything they say will be useful or truthful, but we are already capable of adjusting to that.

This has always been my conjecture. Hallucinating is more like an output that occurs despite the absence of an input.

Confabulating is more like you have 1 or more inputs (drugs, stress, self-interest,trying to save face, your attitude towards truth, etc) and they all go into a homomorphic black box where they interact and bump into each other.

And at the end of the process there is a sociological hash function that puts it all together and manifests it through actions andthe narrative(s) we tell ourselves + others.

The output/behavior/self-talk and reporting are the result of all this and they don't come from nowhere like a hallucination does (in the sense of there being an objective lack of external sensory input, not the neurological/etiological basis which clearly exists).

They have a basis upon which they fundamentally make sense and demonstrate reference to the world around based on how their actions and words are parsed and its not useful to describe LLMs or humans in this way (`hallucinating`) when they are almost always simply in a MonkeyClaw type situation. They have the input and this is a natural result of such input, like any computer does when you process anything else. It follows its instructions and the logic is always inherent to the context and inputs + limitations at play.

Confabulation also often involves "filling in the gaps" one does not have the knowledge or influence or abillity/desire to address so they throw something else that might stick if they luck into avoiding further scrutiny. Sometimes people are "just" lying to you but more often, they are trying to avoid further pain and avarice (maybe even try to gain something) and they simply tell everyone (including themselves) what they want to hear.

Kids are like this, hilarious confabulators. Not hallucinators.

I always took exception to the word “hallucination” with respect to LLMs anyway - they aren’t databases of facts, they’re language models, trained to be really, really good mimics of human communication.

As such, while they are trained on an existing corpus that no doubt contains many facts and many falsehoods, their task is to produce more material that looks like this corpus, not be a source of truth.

So an LLM “hallucinating” or confabulating is not some aberrant behaviour, it’s fundamental to the nature of the beast.

It always worries me when people say they’re using an LLM as a search engine or source of knowledge.

Fun fact, it is believed that Andrej Karpathy (OpenAI, ex AI Dir. at Tesla) had coined the word 'hallucination' in an RNN blogpost back in 2015. https://x.com/karpathy/status/1702916988891193460?s=20
Yeahnah confabulate isn't the right word at all. Confabulate is where you talk amongst your selves in an informal meeting. Often shortened to confab.

Confabulation has a psychiatric meaning, but that's secondary.

Hallucination is a much better word, because its something imagined (involuntarily), but clearly bullshit. Ravings, business bullshit, "my dad has a jetski" are all better words/metaphors to use.

Moreover most people know what hallucination is, whereas most people don't know what confabulate is, more over, if they do, its almost certainly not the meaning you want.

I agree it should have been called confabulation but it will be difficult to change. However, it might also not matter in the following sense. One thing that distinguishes humans is that we have a capacity to recognize whether the information comes from external world (is what we perceive as real world) or from our own thoughts (is what we call dreams). I don't think current crop of LLMs can actually have that distinction, architecturally. So whether the incogruencies with reality happen on the input (hallucination) or the output (confabulation) doesn't really matter, because, from the perspective of LLM, the internal model of the world and the reality of the world are the same thing.

I think that's also a reason why it cannot handle incongruencies in its own mental model; our mental capacity is not only to accept the reality as given (i.e. reality takes precedence over mental model, we recognize that something is just our imagination), but also to impose our model to reality (i.e. we determine that sensors are misleading us due to incongruency with our mental model, and override them as faulty data). But unlike LLMs, humans can internally detect when we do one or the other. Humans understand the distinction between pondering an action and making an action, LLMs don't. LLMs can detect something is amiss (and apologize) but they can't resolve it, because reality and model are the same for them.

“What are LLMs but humans with extreme amnesia and no central coherence?”

LLMs are nothing alike to compare them with anything close to what’s going on in a brain. It’s all probabilistic and the closest comparison I can think of is comparing a bird with a plane - both fly, but well…

There's a somewhat famous movie in Spain called the Miracle of P Tinto[0]. It's a very surreal humor kind of movie, if you fancy that.

The plot starts with the star character who wants to form a family, but when he asked adults how to have kids, they would all pull up the suspenders and sing "tralari, tralari" to evade the question.

And so the poor guy did exactly what he was told: pull his suspenders and sing, for years, to no avail.

I think LLMs are a bit like that: because they can't escape the world of text, they can't go beyond the reality that text is trying to be isomorphic to. Which means that LLMs are forced to infer answers in the context of its training data.

If LLMs were told that in order to have kids you, well, you know, right? Bees and Love? Then LLMs, asked about sex, would infer that the answer has something to do with apiculture and flowers.

[0]: https://en.wikipedia.org/wiki/The_Miracle_of_P._Tinto

I couldn’t help but ask Claude 2 for some better words to name this phenomenon. The response was fascinating:

“Here are some alternative words or concepts that could potentially provide a more nuanced or accurate description of the phenomenon compared to solely using "confabulation":

- Imputation: Generating explanatory responses by imputing or attributing plausible details, characteristics, or causal relationships even when factually incorrect.

- Speculation: The responses reflect a type of informed speculation when direct factual knowledge is lacking, as opposed to confabulation which often implies a pathological or impaired process.

- Heuristic reasoning: The behavior emerges from the model's trained heuristic techniques for reasoning and responding rather than any psychological process like confabulation per se.

- Plausible surmise: The responses take the form of plausible but unfounded surmises put forth due to the system design and limitations rather than an intent to confabulate.

- Bridging inference: The model draws bridging or gap-filling inferences to connect related concepts or complete a response, even if factually ungrounded.

- Suppositional fabrication: The language hints more strongly that the responses are fabricated suppositions not intended to deliberately deceive but still warranting cautious interpretation.

- Educated approximation: The responses could be viewed as a kind of educated approximation or speculation given training-imposed response pressures rather than confabulation due to impairment.

Designating it as solely "confabulation" may over-attribute psychological realism without addressing countervailing technical factors.”

Personally, I thought “Suppositional fabrication” came closest. Maybe just fabrication would work.

I don't think it really matters what we call the phenomenon – whatever they're doing, they're not humans, so whatever human analogy we choose to call it is necessarily imprecise. What does matter is that everyone has a shared vocabulary for what to call hallucination or whatever.
The Magic 8 Ball (invented in 1946) also confabulate not hallucinate. It is just a evergreen hype to attach human properties (in a psychiatric context... brain damage... amnesiac patient...) to objects that have zero connection to human intellect.

IBM Watson from 2010 won Jeopardy. Markov Chains were described in 1906 and used for a long time to simulate human activities (including simulation of blabbering of mentally handicapped). In Ancient Rome there were "talking boards" (some kind of Ouija boards) to predict the successor to the king. It is just a tendency for perception to do such associations, similar to pareidolia.

I asked Bard to summarize this by pasting the URL.

It made something up!! (it said the author talked about a poem about cat and mice as an example of confabulation)

Then I asked "are u confabulating? because i do not see that quote in the original article"

  You are correct. I confabulated that information. I apologize for the mistake.I am still under development, and I am always learning. I will try my best to avoid making similar mistakes in the future.Thank you for pointing out my error. It helps me to improve.
Anthropomorphizing LLM is at best misleading, but probably extremely dangerous. At first I didn’t care too much, the “hallucination” term got traction, now the ladies on The View are saying it and thinking they “understand ai”.

The public needs to know how LLM work, they don’t need bs abstractions that blur or even erase the lines between humans and technology. LLM don’t hallucinate, they don’t confabulate. They are machines that operate according to principles all educated people need to understand. It’s the duty of the people who understand it better to educate them, not with half truths and metaphors (though I sort of like rms’ “bullshit generator” - it’s obviously wrong enough/funny in a way that lets people know it’s not a precise technical definition, while coming very close to being an accurate technical definition.)

The danger of ai is this: confusing human and machine. They are separate and different, and must remain that way. LLM aren’t capable of hallucinating, confabulating any more than they are capable of falling in love or serving on jury duty. They repeat the patterns found in the text you feed them. That’s it.

Why call it anything other than an error or bug? When my server spits out data I didn’t expect or want, I don’t say it hallucinated or confabulated.

When the hurricane model gets the landfall site wrong, what do we call that?

There was just a HN thread on police crime models. When those are wrong are they “hallucinating” crime?

Why personify these statistical models at all?

A language model is just that.. it predicts probabilities of words given some prefix. Of course such a thing cannot hallucinate. What is confabulating is the human imposed process of selecting a random word weighted on the probability distribution returned. Obviously it's going to cause a problem. Imagine communicating with someone and they ask a yes or no question and you attempt to tell them the truth which is that you are 70 percent yes and 30 percent no. And then they apply a randomized picking algorithm that weights those responses and when it comes back as 'no', they go and tell everyone you lied.

Unless we make word output an actual part of LMs we can't see they are lying or not lying. They only produce probabilities of word outputs based on the context.

Call it as you wish, it's still enervating when you ask it the same poetry three times in a row, and it invents a new poetry each time, and you force it to recognize that it doesn't match the previous hallucinations/confabulations, and it apologizes and proceeds to a new artistic spree.
Regardless of the accuracy of the origin story, Confabulation Theory[0] is the first place that I can recall this sort of explanation that combines knowledge with "honest lies". I don't know how accepted his definition of the word confabulation really is, but it seems a good descriptor of what's going on. Not a hallucination, but more an issue in walking the connections between concepts, with knowledge being stored in the edges.

[0]: http://www.scholarpedia.org/article/Confabulation_theory_(co...

The question is, why don’t LLMs admit when they don’t know an answer?

If they can’t, how do we make them so they can?

Plenty of humans seem to have no problem admitting when they don’t know something.

The term of “hallucination” referred to generative ML models have been used in the literature for years, at least since GANs.L
LLMs answer what unfoundly self-convinend people answer when you ask them for diretion on the street. Convincing, not necessarily correct. They are trained on the “data highway“ mostly. A highway full of that type of people, so that behaviour was to be expected.
sure... "AI's are hallucinating!"

Maybe stop calling probalistic text generators "intelligence" and then you wont see "hallucinations", but simply "garbage" output?

Both terms are problematic as they anthropomorphize an algorithm.
LLMs answer with unwarranted confidence.

LLMs speak with confidence about disciplines they don’t actually understand, like sociology and economics.

So we’ve finally automated Hacker News comments!

Hallucinate means false perception which doesn't describe what LMs do when they're making up plausible sounding things
Do most people understand exactly what confabulate means?

I'd rather just say that they're bullshitting. Everyone knows what it means, and it has a precise academic definition[0] within the field of bullshitology.

[0] https://philosophynow.org/issues/53/On_Bullshit_by_Harry_Fra...

Surely the one trained on Erowid data is hallucinating?
They don't hallucinate, they just get things wrong.
> this isn’t actually correct terminology

This entire post could have been reduced to this (reasonably correct I think) claim.

> if we recognize that what LLMs are really doing is confabulating, we can try to compare and contrast their behaviour with that of humans

Except that the goal for the above is hopelessly misguided. If your focus in making any useful claim about LLMs or RNNs/GANs more broadly is how similar they are to human behavior you're already way off the path.

tl;dr reasonable claim, reasonable conclusion, but entirely indefensible and superstitious motivation