1) Hallucinations often appear because LLMs are designed to create fluent, coherent text.
2) LLMs have no understanding of the underlying reality that language describes.
3) LLMs use statistics to generate language that is grammatically and semantically correct within the context of the prompt.
It sacrifices accuracy for being good at conversations as it is designed to do. All these criticisms of hallucinations are missing the point.
Generative AI is generative and basically is a specialist at making things up. It’s going to take things like multiprompting, network AIs that fact check output, and a host of other technologies or even entirely new models of AI to solve these problems, but don’t make the mistake of thinking that the system is supposed to be working without hallucinations right now — that’s not what it’s optimized for.
What they just don't care about is communicating being out of distribution or making whack predictions. This becomes a problem for humans because they're perfectly fine making things up when the above fail.
But by all accounts, they do learn to distinguish these things. The computation is very much aware when it is going way off base.
GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - https://arxiv.org/abs/2305.14975
Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334
Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets - https://arxiv.org/abs/2310.06824
LLMs absolutely have concepts that extend beyond just words. As early as 2008 (back when there were no large language models, only "regular" language models), we've been able to demonstrate things that seem to me like the model is learning abstract concepts. For a classic example, see Linguistic Regularities in Continuous Space Word Representations [0], a 2013 paper that talks about a model with a latent space where one can take the vector representations of the words "king", "queen", "man", and "woman" and literally perform the arithmetic `king - (man - woman) ≈ queen`. To me, this clearly demonstrates that the model "understands" the concept of gender, represented by a vector in the model's latent space. The set of numbers that you get from the subtraction `man - woman` represent the concept of the difference between the male and female genders, without needing a specific word tied to that representation of the concept.
It's debatable whether or not that counts as "understanding", but that's more of a semantic debate about what it means to understand, not a debate about the model's internal knowledge of the world and ability to do things with that knowledge.
I always find this point a bit odd because humans aren't "optimized" to be correct either.
> "LLMs have no understanding of the underlying reality"
I struggle with this one because I see both sides of it. I was making a prompt the other day and gave a CSV file as an input and told the LLM if could only answer with values from one column and it did exactly as I asked. It's hard for me to see things like that and not believe it has an understanding at some level.
Why we baby AI on this front, I have no idea.
Not if people need to be reminded that, as you say, LLMs are not designed to give reliable answers. Many, many people appear to believe that they are, since they rely on the answers in all kinds of contexts. For example, in past HN threads on LLMs, I have seen people say that they rely on LLM-generated code.
Not true. They are slowly gaining an understanding of reality by reverse engineering the relationships built into human languages. The only reason LLMs are getting better is because they are better modeling the world. At some point the only way to improve token prediction is to gain an understanding of the world.
The essence of bullshit (as he defines it) is that it is neither truth nor lie but rather unconcerned with veracity altogether. A bullshitter wants to appear smart, or convince others to agree with them, or some other conversational objective. Bullshit statements may or may not ultimately turn out to be true, but the thing that makes them bullshit is that the bullshitter didn't actually know when they said it.
Hence I propose a far more succinct term to describe the phenomena at hand. A far more accurate terminology than the presently vogue "hallucinating": AB. Artificial Bullshit.
If instead it told the story of someone short and obese, with details on how he got to play basketball, that would be creative. Of course, you can ask an AI to tell the story of a short and obese basketball player, and it will tell you a likely story for such a player, but again that's not creativity, it is just the AI filling the blanks based on whatever similar stories it had in its training set.
> Meet Ethel, a 4'9" grandmother of six, with a penchant for knitting and a mastery of Sudoku. With bifocal glasses perched on her nose, she's far from your typical basketball star. But what Ethel lacks in height and athleticism, she makes up for with an uncanny sixth sense for predicting opponents' moves and an underhand free-throw that rivals the best in the league. With her floral-print headband and orthopedic shoes squeaking down the court, she's both an anomaly and a secret weapon on her community center's basketball team.
https://chat.openai.com/share/62180301-b7ae-46fe-bd15-bd6973...
> Jack "Shadow" Carter was a prodigy dismissed for his short stature, standing only 5'7". Ignored by scouts and overshadowed by taller players in high school, he developed a unique playing style that exploited his low center of gravity and agility. He became a master of steals and assists, zipping around the court like a shadow, hence his nickname.
(Though I get your point that "hallucinations" will tend towards lowest-common-denominator answers, not creative answers.)
The article says "well, sometimes what it makes is bad".
Well big deal. A lot of human-created art is awful too.
Yeah but who wants to consume art purely generated by AI (that is, not human-created with AI support)? Most art sites have had blanket bans, or at least required tagging, on ai-generated art because people hate it so much.
Or to put it another way: why are you in the comment section of Hacker News, and not just asking ChatGPT to generate social media comments on the article?
That's how lots of science, innovation, & learning work: generate many superficially-plausible candidates via a fast-and-loose process, then refine with a more rigorous evaluation.
That AIs, in the form of LLMs, are now doing this so well was unexpected, and progress in checking 'hallucinations' is proceeding very fast.
(Fortunately, the article is less dismissive than the headline, recognizing these model's potential & mainly urging an understanding of the limitations.)
> It might be better to say that everything GPT does is a hallucination, since a state of non-hallucination, of checking the validity of something against some external perception, is absent from these models.
I try to explain this to people who are obsessed with using ChatGPT to tell them things. So far I've been telling them something like: "it does not attempt to provide you valid information, it's optimizing for what would read like a reasonable continuation of the conversation, which is really not the same thing."
The AI doesn't know something so it just invents something. We usually call that "bullshitting", or in more polite crowds, "lying".
When a human thinks it’s able to determine the difference between imagination and fact, but when a human reads words on a screen it isn’t?
Drivel.
> Unfortunately, this promotes a misunderstanding of how large language models (LLMs) work... It might be better to say that everything GPT does is a hallucination, since a state of non-hallucination, of checking the validity of something against some external perception, is absent from these models.
Let me turn it around:
> Unfortunately, this promotes a misunderstanding of how brains work... It might be better to say that every question a human answers without research is a hallucination, since a state of non-hallucination, of checking the validity of something against some external perception, is absent when simply answering a question.
Obviously nonsense. If you're going to write an article about misconceptions you'd better make sure you are right!
(Though it is silly to celebrate hallucinations; they're definitely not desirable.)
I'm fine with the term "hallucination," though. Hallucinations are, by definition, not real. The term emphasizes a detachment from reality that LLMs possess, and the general public is just beginning to grasp.
It is also sensationalist, implying we don't really know what's going on, why is this LLM hallucinating?
It even sounds like an excuse: Hey my LLM did not come out with reasonable answer, but that is only because it was hallucinating, just like humans sometimes do, so it is even more human-like than we thought. See. Or may it was drunk! That explains it.
No, it's just that LLMs are sometimes on the topic, sometimes not. When somebody says they are "hallucinating" it does not mean they are working in some kind of extra-ordinary mode of operation. They are working just as usual.
What explains what some people (want to) call LLM "hallucinating" is that LLMs sometimes make sense, sometimes they don't.
So "irl", we see people like Alex Jones that get up on their big platforms and start spewing nonsense, but if they sound confident enough and it confirms what you want out of the world, then people latch onto it as fact and don't bother to verify. You see this across the internet. Just today I saw a story on instagram that had been re-posted and the person that was talking about it was many degrees removed from the original, but believed it to be real. Looking through the comments, I had to scroll past 50+ comments to find someone who finally called it out as fake. Everyone else was just posting "no way", "wow, I never knew". You never knew because its completely made up. But when we hear someone speak with confidence and we don't care enough to fact-check, then people just believe it.
This is no different than AI. AI models sound confident and reliable. We assume they are making proclamations based on fact but they aren't always (or "usually" in my experience). Many people blindly believe the AI models because they sound reliable and confident in the way they speak. They never say "i don't know".
What AI is doing is problematic for sure. On one hand I want to take the pitchforks and revolt. But on the other hand I look around and realize, that even if we fixed it or vanquished this enemy, we still have a bunch of talking heads doing the same thing.
Maybe AI hallucinations are actually the most human element of AI.
Truth is hard. Maybe too hard for a mere machine. But dramatic narrative and quirky dialogue might be quite doable.
3000 chapter litrpg fantasy generated overnight.
On the (rare) occasion I find it useful to avoid this very normal tendency I ask myself if it would make sense to apply the same framing to the output of an AI image generator.
https://chat.openai.com/share/e213e0bd-2838-45e2-9942-e52954...
https://chat.openai.com/share/64bc62d3-042c-40c4-8d50-8e28ce...
Hilariously, plugging the example in the article into Bing enhanced ChatGPT-4, ChatGPT-4 w/ Bing hallucinates because of that very article!
https://chat.openai.com/share/4fe50933-8436-44ad-a778-6297ca...
If you tell ChatGPT-4 w/ Bing to ignore thereader.mitpress.mit.edu where the article is hosted, it doesn't hallucinate the string "Evolution by Any Other Name?".
https://chat.openai.com/share/b8a94a8e-723f-40b6-b7c5-698dcd...
source: https://chat.openai.com/share/4fe50933-8436-44ad-a778-6297ca...
Hallucination is always and everywhere used as a negative term for LLMs. And it is seen as a problem/challenge that we should get rid of.
Did not read the article after seeing such a wrong title.
Given that these AI systems just like any other machine operate at scale, automated, and fast, they must be precise and transparent, that is where the work should be.
I've only heard of people trying to "solve" them.
Nor are they capable of creativity.