She can’t read, write, handle money, but confabulates engaging and convincing stories about receiving mail, haggling with her bank, her boyfriend. She speaks with proper grammar, intonation, etc, but doesn’t demonstrate any understanding of what she talks about beyond the words. She doesn’t have a bank account or a boyfriend.
Trying to find more and I can’t find anything about her on the internet, I would have to look at the references. Maybe it’s questionable.
Anyway, I like to think about what the hell we even mean when we say we “know something”, and where we should actually put the bar for AI. Where is the line between confabulation and knowledge? Is knowledge just confabulation that happens to be correct?
I feel like it’s possible that Denys could feel that she “knows” what she is saying, but is simply trapped in, with a working knowledge of the world and is just unable to translate it into actions, and words are just her only degree of freedom.
I don’t remember where I was going with this, but I agree that “confabulation” is a much better term to use, but we also shouldn’t discount what it means to be able to confabulate.
The tactic is simply to talk relevantly until the listener fools themselves into thinking they understand you, except that "understanding" will be an illusion and 100% manager's own analytical effort. And it's honestly embarrassing how often and well that works. Quite a few of those guys were promoted first, too.
I'm pretty sure our upcoming future with AI will be like that quote from forest ranger at Yosemite National Park, on why it is hard to design the perfect garbage bin to keep bears from breaking into it: "There is a considerable overlap between the intelligence of the smartest bears and the dumbest tourists."
The whole idea of a criteria to separate humans from AI is a misconception. I think any human who confuses smooth talking with making sense will be ruthlessly exploited in relatively near future. There are humans who already do this successfully to other humans, what do you think will happen when they get to automate their efforts.
Instead I would say they DREAM. Dreams follow some logic but basically they are disjoint from reality. They involve same characters as our real-life experiences but what those characters do in a dream is not based on reality.
So one could think that ALL LLMs ever do is dream but much of the time their dreams are dreams which feel very real.
Our dreams are basically an LLM based on our real experiences but recombined in arbitrary ways, still retaining some logical structure but now detached from reality because when you recombine different experiences including our thoughts from real life, the result cannot really reflect reality very well. LLMs are better in this respect than our dreams, but sometimes their "dreams" get really obviously detached from reality.
Or if you want to use the term "hallucinate" with LLMs then fine but its more like they hallucinate all the time, it's just that sometimes, even often, their hallucinations seem to agree with the reality so well that we cannot tell the difference.
LLMs do not describe reality because they cannot experience reality, all they can give us is an "average description" created from many existing descriptions.
One way of looking at it is to say that all LLMs tell us is hearsay.
GPT-4 logits calibration pre RLHF - [https://imgur.com/a/3gYel9r](https://imgur.com/a/3gYel9r)
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - [https://arxiv.org/abs/2305.14975](https://arxiv.org/abs/2305...
Teaching Models to Express Their Uncertainty in Words - [https://arxiv.org/abs/2205.14334](https://arxiv.org/abs/2205...
Language Models (Mostly) Know What They Know - [https://arxiv.org/abs/2207.05221](https://arxiv.org/abs/2207...
It seems very much like hallucinations aren't a randomness or representation issue. The computation already knows. It just doesn't care about telling you this
So not unlike early LLM chat in areas where the model had not a lot of data about.
Here is another interview, this one between a fourteen-year-old girl called Denyse and the late psycholinguist Richard Cromer; the interview was transcribed and analyzed by Cromer's colleague Sigrid Lipka.
> I like opening cards. I had a pile of post this morning and not one of them was a Christmas card. A bank statement I got this morning!
< [A bank statement? I hope it was good news.]
> No it wasn't good news.
< [Sounds like mine.]
> I hate . . . , My mum works over at the, over on the ward and she said "not another bank statement." I said "it's the second one in two days." And she said "Do you want me to go to the bank for you at lunchtime?" and I went "No, I'll go this time and explain it myself." I tell you what, my bank are awful. They've lost my bank book, you see, and I can't find it anywhere. I belong to the TSB Bank and I'm thinking of changing my bank 'cause they're so awful. They keep, they keep losing . . . [someone comes in to bring some tea] Oh, isn't that nice.
< [Uhm. Very good.]
> They've got the habit of doing that. They lose, they've lost my bank book twice, in a month, and I think I'll scream. My mum went yesterday to the bank for me. She said "They've lost your bank book again." I went "Can I scream?" and I went, she went "Yes, go on." So I hollered. But it is annoying when they do things like that. TSB, Trustees aren't. . . uh the best ones to be with actually. They're hopeless.
I have seen Denyse on videotape, and she comes across as a loquacious, sophisticated conversationalist—all the more so, to American ears, because of her refined British accent. (My bank are awful, by the way, is grammatical in British, though not American, English.) It comes as a surprise to learn that the events she relates so earnestly are figments of her imagination. Denyse has no bank account, so she could not have received any statements in the mail, nor could her bank have lost her bankbook. Though she would talk about a joint bank account she shared with her boyfriend, she had no boyfriend, and obviously had only the most tenuous grasp of the concept "joint bank account" because she complained about the boyfriend taking money out of her side of the account. In other conversations Denyse would engage her listeners with lively tales about the wedding of her sister, her holiday in Scotland with a boy named Danny, and a happy airport reunion with a long-estranged father. But Denyse's sister is unmarried, Denyse has never been to Scotland, she does not know anyone named Danny, and her father has never been away for any length of time. In fact, Denyse is severely retarded. She never learned to read or write and cannot handle money or any of the other demands of everyday functioning.
Denyse was born with spina bifida ("split spine"), a malformation of the vertebrae that leaves the spinal cord unprotected. Spina bifida often results in hydrocephalus, an increase in pressure in the cerebrospinal fluid filling the ventricles (large cavities) of the brain, distending the brain from within. For reasons no one understands, hydrocephalic children occasionally end up like Denyse, significantly retarded but with unimpaired—indeed, overdeveloped—language skills. (Perhaps the ballooning ventricles crush much of the brain tissue necessary for everyday intelligence but leave intact some other portions that can develop language circuitry.) The various technical terms for the condition include "cocktail party conversation," "chatterbox syndrome," and "blathering."
Source: http://f.javier.io/rep/books/The-Language-Instinct-How-the-M...For instance, before transformers, when we finetuned ESRGAN for specific styles/functions, we used to say it would "hallucinate" in detail that couldn't possibly be in the original image pixels. GenAI wasn't much of a thing back then, and we never thought of the small GAN models as "remembering" things even though thats common language in transformer models. It felt more like we were teaching the GAN a job or style, not encoding memory.
And I think it fit! "Hallucinating" detail into a blurry pixel blob feels like a more accurate analogy to the human condition, and we weren't at the point where it was "confabulating" and generating big objects out of the blue like a diffusion model can.
See https://blog.research.google/2015/06/inceptionism-going-deep... and https://distill.pub/2017/feature-visualization/.
I would differentiate these dreamy generations (i.e., things that we don't see ordinarily) from generations that are incorrect. Sure, there is overlap.
https://neurosciencenews.com/peripheral-vision-brain-illusio...
Both words imply that there's something erroneous or anomalous happening when an LLM emits incorrect facts, as if the model can be in either a correct or incorrect state. But from a system perspective, there's no difference between the two scenarios: the model is consistently generating the most probable output based on its weights and context. It doesn't sometimes confabulate... it always confabulates.
The concept of "truth" or "falsehood" are something purely external to the model that humans bring when evaluating the output. They aren't part of the model itself.
I recall seeing a post on here that roughly just said, "stop saying hallucinate", which sure, is technically correct ("the best kind of correct"), but not super helpful. Contrast that with this post, which offers a constructive suggestion: replace it with "confabulate".
I agree that C is an improvement over H, for different reasons from the author. Here's why: From a public comms standpoint, the big problem with saying "hallucinate" is that it leads to misconceptions, because people think they know what "hallucinate" means, and may anthropomorphize their pop psych grasp of H onto the LLM.
Just like H, "C" seems to have a technical meaning in the psych literature, but it's far less in common parlance than "H", so lay readers have no fixed idea of what exactly "C" means. Nobody goes around experimenting with "confabulagens" to seek a spiritual high. Also, "confabulate" retains a non-clinical meaning that's similarly obscure, from long before the psychologists adopted it, and it's pretty evocative of what's really going on here (e.g. via related words like "fable" and "fib"). So "confabulate" seems ripe for adoption to describe this specific computational phenomenon, as we've done to so many other mundane words in the past (deprecate, and so forth).
If the psychiatric "C" phenomenon is also a better analogy for LLM behavior, as OP claims (and some posters dispute?), so much the better. But my main reason for liking C, is that it's less likely than H to make people unduly confident that they understand it.
(But even if the community came to a consensus at this late stage that "C" is a better term, the subsequent problem of updating every stale paper and post that uses the old term would run us up against another doozy: cache invalidation.)
Hallucinate generally implies experiencing sensory inputs which aren't there, but having a rational reaction to them - in the popular psyche though, people who hallucinate are still "crazy" - unfairly so, because you can function normally with hallucinations under quite a number of circumstances (i.e. I remember someone saying that they realized that if they saw people without faces, it was fine because they weren't really there - apparently quite common, since face recognition is a different part of the brain).
Basically: reacting to invalid sensory inputs implies rational behavior, but to non-real inputs.
Confabulation on the other hand is different because it's closer to "trying to rationalize an autonomous behavior without knowledge of the sensory inputs causing it" (maybe, it's a bizarre phenomenon). In split-brain patients in lab settings, it manifests as being unable to answer a question about why you're having a physiological response to imagery which is being shown only to the right-side of the brain, whereas language is processed on the left - but rather then be confused, people will apparently make something up that they're unaware is actually a "lie" (in quotes because, well, they're not lying - as far as we know there's no intent or even knowledge that is a lie).
Where this leaves us with LLMs I don't really know, because it feels like an imperfect descriptor: except perhaps with the context that LLMs are fairly limited networks, and so to some extent the whole "being wrong" phenonmenon is in fact just a failure of the attention mechanism to be able to draw the right data together.
Indeed, the Latin verb "confabulor has the meaning of "discuss, converse" [0], not specifically "inventing things", and Romance languages already have words derived in sound and meaning from this classic Latin sense. For instance French already has "confabuler" meaning "speak familiarly with someone" [1]. Thus in Romance languages, this proposed usage of "confabulate" imported from the psychiatric word could be understood poorly due to preexisting similar-sounding words that do not include the notion of "making things up".
Just proposing thoughts here, not advocating for a final decision. Maybe it's still good to adopt this wording!
[0] : https://en.m.wiktionary.org/wiki/confabulor#Latin [1] : https://www.dictionnaire-academie.fr/article/A9C3479
The answer to "What are LLMs but humans with extreme amnesia and no central coherence?" is that they are not like humans at all, and only resemble them superficially due to our anthropomorphic tendency to infer a mind upon seeing fluent and roughly correspondent text.
But neurotypical, neurodiverse, perfectly functional people, etc. all confabulate or do something similar on a regular basis, in verbal and written mediums, and often do so in good faith. It's human instinct to communicate, even if you are uncertain and unaware of the full context of the discussion.
Teachers, customer service reps, executives, shop keepers, doctors, nurses, domain experts, authors of textbooks, it doesn't matter who it is, they'll probably confabulate or equivocate or do some other type of communication that isn't immediately useful. Yet it's still a useful activity to just talk to someone or read a less than rigorous book for the purposes of learning (discounting the relationship forming part, which is also useful). And so is using LLMs, even for casual users. So long as they understand that limitation, whether its with a chatbot or a real person. Not everything they say will be useful or truthful, but we are already capable of adjusting to that.
Confabulating is more like you have 1 or more inputs (drugs, stress, self-interest,trying to save face, your attitude towards truth, etc) and they all go into a homomorphic black box where they interact and bump into each other.
And at the end of the process there is a sociological hash function that puts it all together and manifests it through actions andthe narrative(s) we tell ourselves + others.
The output/behavior/self-talk and reporting are the result of all this and they don't come from nowhere like a hallucination does (in the sense of there being an objective lack of external sensory input, not the neurological/etiological basis which clearly exists).
They have a basis upon which they fundamentally make sense and demonstrate reference to the world around based on how their actions and words are parsed and its not useful to describe LLMs or humans in this way (`hallucinating`) when they are almost always simply in a MonkeyClaw type situation. They have the input and this is a natural result of such input, like any computer does when you process anything else. It follows its instructions and the logic is always inherent to the context and inputs + limitations at play.
Confabulation also often involves "filling in the gaps" one does not have the knowledge or influence or abillity/desire to address so they throw something else that might stick if they luck into avoiding further scrutiny. Sometimes people are "just" lying to you but more often, they are trying to avoid further pain and avarice (maybe even try to gain something) and they simply tell everyone (including themselves) what they want to hear.
Kids are like this, hilarious confabulators. Not hallucinators.
As such, while they are trained on an existing corpus that no doubt contains many facts and many falsehoods, their task is to produce more material that looks like this corpus, not be a source of truth.
So an LLM “hallucinating” or confabulating is not some aberrant behaviour, it’s fundamental to the nature of the beast.
It always worries me when people say they’re using an LLM as a search engine or source of knowledge.
Confabulation has a psychiatric meaning, but that's secondary.
Hallucination is a much better word, because its something imagined (involuntarily), but clearly bullshit. Ravings, business bullshit, "my dad has a jetski" are all better words/metaphors to use.
Moreover most people know what hallucination is, whereas most people don't know what confabulate is, more over, if they do, its almost certainly not the meaning you want.
I think that's also a reason why it cannot handle incongruencies in its own mental model; our mental capacity is not only to accept the reality as given (i.e. reality takes precedence over mental model, we recognize that something is just our imagination), but also to impose our model to reality (i.e. we determine that sensors are misleading us due to incongruency with our mental model, and override them as faulty data). But unlike LLMs, humans can internally detect when we do one or the other. Humans understand the distinction between pondering an action and making an action, LLMs don't. LLMs can detect something is amiss (and apologize) but they can't resolve it, because reality and model are the same for them.
LLMs are nothing alike to compare them with anything close to what’s going on in a brain. It’s all probabilistic and the closest comparison I can think of is comparing a bird with a plane - both fly, but well…
The plot starts with the star character who wants to form a family, but when he asked adults how to have kids, they would all pull up the suspenders and sing "tralari, tralari" to evade the question.
And so the poor guy did exactly what he was told: pull his suspenders and sing, for years, to no avail.
I think LLMs are a bit like that: because they can't escape the world of text, they can't go beyond the reality that text is trying to be isomorphic to. Which means that LLMs are forced to infer answers in the context of its training data.
If LLMs were told that in order to have kids you, well, you know, right? Bees and Love? Then LLMs, asked about sex, would infer that the answer has something to do with apiculture and flowers.
“Here are some alternative words or concepts that could potentially provide a more nuanced or accurate description of the phenomenon compared to solely using "confabulation":
- Imputation: Generating explanatory responses by imputing or attributing plausible details, characteristics, or causal relationships even when factually incorrect.
- Speculation: The responses reflect a type of informed speculation when direct factual knowledge is lacking, as opposed to confabulation which often implies a pathological or impaired process.
- Heuristic reasoning: The behavior emerges from the model's trained heuristic techniques for reasoning and responding rather than any psychological process like confabulation per se.
- Plausible surmise: The responses take the form of plausible but unfounded surmises put forth due to the system design and limitations rather than an intent to confabulate.
- Bridging inference: The model draws bridging or gap-filling inferences to connect related concepts or complete a response, even if factually ungrounded.
- Suppositional fabrication: The language hints more strongly that the responses are fabricated suppositions not intended to deliberately deceive but still warranting cautious interpretation.
- Educated approximation: The responses could be viewed as a kind of educated approximation or speculation given training-imposed response pressures rather than confabulation due to impairment.
Designating it as solely "confabulation" may over-attribute psychological realism without addressing countervailing technical factors.”
Personally, I thought “Suppositional fabrication” came closest. Maybe just fabrication would work.
IBM Watson from 2010 won Jeopardy. Markov Chains were described in 1906 and used for a long time to simulate human activities (including simulation of blabbering of mentally handicapped). In Ancient Rome there were "talking boards" (some kind of Ouija boards) to predict the successor to the king. It is just a tendency for perception to do such associations, similar to pareidolia.
It made something up!! (it said the author talked about a poem about cat and mice as an example of confabulation)
Then I asked "are u confabulating? because i do not see that quote in the original article"
You are correct. I confabulated that information. I apologize for the mistake.I am still under development, and I am always learning. I will try my best to avoid making similar mistakes in the future.Thank you for pointing out my error. It helps me to improve.The public needs to know how LLM work, they don’t need bs abstractions that blur or even erase the lines between humans and technology. LLM don’t hallucinate, they don’t confabulate. They are machines that operate according to principles all educated people need to understand. It’s the duty of the people who understand it better to educate them, not with half truths and metaphors (though I sort of like rms’ “bullshit generator” - it’s obviously wrong enough/funny in a way that lets people know it’s not a precise technical definition, while coming very close to being an accurate technical definition.)
The danger of ai is this: confusing human and machine. They are separate and different, and must remain that way. LLM aren’t capable of hallucinating, confabulating any more than they are capable of falling in love or serving on jury duty. They repeat the patterns found in the text you feed them. That’s it.
When the hurricane model gets the landfall site wrong, what do we call that?
There was just a HN thread on police crime models. When those are wrong are they “hallucinating” crime?
Why personify these statistical models at all?
Unless we make word output an actual part of LMs we can't see they are lying or not lying. They only produce probabilities of word outputs based on the context.
[0]: http://www.scholarpedia.org/article/Confabulation_theory_(co...
If they can’t, how do we make them so they can?
Plenty of humans seem to have no problem admitting when they don’t know something.
Maybe stop calling probalistic text generators "intelligence" and then you wont see "hallucinations", but simply "garbage" output?
LLMs speak with confidence about disciplines they don’t actually understand, like sociology and economics.
So we’ve finally automated Hacker News comments!
I'd rather just say that they're bullshitting. Everyone knows what it means, and it has a precise academic definition[0] within the field of bullshitology.
[0] https://philosophynow.org/issues/53/On_Bullshit_by_Harry_Fra...
This entire post could have been reduced to this (reasonably correct I think) claim.
> if we recognize that what LLMs are really doing is confabulating, we can try to compare and contrast their behaviour with that of humans
Except that the goal for the above is hopelessly misguided. If your focus in making any useful claim about LLMs or RNNs/GANs more broadly is how similar they are to human behavior you're already way off the path.
tl;dr reasonable claim, reasonable conclusion, but entirely indefensible and superstitious motivation