> The nice thing about science is that it routinely features “error bars” on its graphs, showing both the finding and the degree of confidence in its accuracy.
> AI/ML products in general don’t have them.
> I don’t see how it’s sane or safe to rely on a technology that doesn’t have error bars.
Exactly, there is no news here. Foundation LLMs are generative models trained to produce text similar to the text on which they were trained (or for the pedantic: minimize perplexity).
They are not trained to output facts or truths or any other specific kind of text (or for the pedantic: instruction-tuned / rlhf types are trained to produce text that humans like after they are trained to minimize perplexity).
"Hallucination" is a terrible term to apply to LLMs. LLMs ONLY produce hallucinations.
Yet even here on HN it isn't uncommon to see top/high ranked comments suggesting baby AGI, intelligence, reasoning, and all that. I suspect I'll get replies about how GPT can do reasoning and world models. Hell, I've seen very prominent people in the space make these and other obtuse claims. A lot of people buy into the hype and have a difficult time distinguishing all the utility from the hype, not understanding that attacking the hype is not attacking utility (plenty of things are overhyped but useful).
I keep saying ML is like we've produced REALLY good chocolate, then decided that this was not good enough so threw some shit on top and called it a cherry. The chocolate is good enough, why are we accepting a world where we keep putting shit on top of good things? We do realize that at some point people get more upset because there's more shit than chocolate, right? Or that the shit's taste overpowers the chocolate's (when people are talking about pure shit, it's because they've reached this point). It's an unsustainable system that undermines all chocolate makers and can get chocolate making banned. For what? Some short term gains?
True, although of course in many contexts facts/truth are the best prediction. Maybe we're just not training them as well as possible.
I've argued that to really fix "hallucinations" (least-worst predictions) these models really need to be aware of the source/trustworthiness of their training data, and would presumably learn that trusted sources better help predict factual answers.
However, it turns out that these models do often already have a good idea of their own confidence levels and whether something is true or not, but don't seem to know when to use that information.
Hallucinations seems to be an area where the foundation companies seem confident they can make significant improvements, although I'm not sure what techniques they expect to (or are) using.
> "Hallucination" is a terrible term to apply to LLMs. LLMs ONLY produce hallucinations.
I don't think that's a great way to characterize it. I prefer the word bullshitting vs hallucinations (they "know" what they are doing), but let's just call it what it is - they statistically predict, and some predictions are better than others. Per human-preference fine tuning, these models are also "trying to please", and I wonder if that has been at least a small part of the problem - predicting a low confidence (when they are aware of it) continuation rather than "I don't know" because human evaluators have indicated a preference for longer and/or more specific answers.
Whether a model is good or bad at recalling facts from its training material matters to end-users and GPT-4, Claude 3 Opus (and bigger models in general?) tend to be better at this than other models.
Rather than hallucinations being treated like an interesting puzzle or paradox at the center of AI intrinsically, I think it's incidental to the types of models that have been trained, and it's conceivable that they could be trained against a notion of reliable sources and the relation between their statements and such sources.
Is the result substantially different from the result coming from biology?
IIRC, a theory is human minds evolved to perceive, store and evaluate not for facts or truths but for evolutionary fit.
What they are weak at is exactly the same stuff we're actually weak at: accurately recalling facts. Except they are able to recall vastly more stuff than any human being would be able to recall. An inhuman amount of facts actually. The core issue is that when asked sufficiently open questions, these models tend to take some liberties with the facts. But most of the knowledge tests that are used to benchmark LLMs, would be hard to pass for the vast majority of humans on this planet as well.
Worse, if you follow the public debate on various political topics a bit you realize that it features a lot of people suffering from confirmation bias parroting each other. Populist politicians seem to get away with a lot of stuff that would put most LLMs to shame; seemingly without affecting their popularity.
IMHO, LLMs by themselves can't be trusted to get things right but paired with some subsystems to produce references, check things, look things up, they become quite capable. Also, it helps asking the right targeted questions and constrain them a little.
We pay for chat gpt at work. At 20$ per month per user, it's pretty much a no-brainer. Do we blindly trust it? Absolutely not. Do we get shit tons of value out of it? Yes. I program with it, I brainstorm with it, I let it review text, I use it to work out bullet points into a coherent narrative, I use it to refine things, I use it to generate unit tests, etc. Is it flawless? No. But it sure saves me a lot of time. And getting the same value from people tends to be a lot harder/more expensive.
This dismissive "it's just a stochastic parrot" type criticism kind of misses the forest for the trees. If you've ever observed toddlers repeating stuff adults tell them, you'd realize that we all start out as stochastic parrots. Forget about getting any coherent/insightfull statement out of a toddler. And most adults aren't that much better and would fail most of the tests we throw at LLMs.
Seems pedantic. If you define "hallucinations" to include correct responses, then nobody cares whether something is a "hallucination," the only important question is -- is it right?
And the answer is yes, across a huge number of metrics [at least for gpt-4]. It'd beat you at jeopardy, it'd beat you at chess, it'd beat you at an AP-biology exam, it'd beat you at a leet code competition.
Enough with semantics.
If you train them on text that is largely true, then the output is also largely true. To ignore this is to miss the point of LLMs altogether.
They are not trained to output facts or truths or any other specific kind of text
but we can just patch the truth in with RLHF! /s
AI-skeptic: "I can come up with AN example where an AI is worse than an encyclopedia, and it's confident! Burn it all down! Hallucinations!"
AI-user: "I literally couldn't care less about that. It's more accurate than any person I know personally on almost every topic and makes code that works 70% of the time in 5 seconds."
I find them tremendously useful but the bullshitting, prudishness, bland mainstream views on everything, etc. can be quite irritating.
You know, being well informed beyond marketing materials.
The user then might demand action from the service provider thereby improving the service.
Instead of you know, yolo.
That's a gross understatement. I'd prefer something like "misinformation presented as truth"
Like, we have the internet. You don’t have to ask people you know personally! You can look stuff up!
And yes, it clearly is worse than an encyclopedia. And we have those!
LLMs should be used in use cases that play to their strenghts: conjuring up stories, summarizing content, polishing text drafts etc. Search is about the worst use case for an LLM.
Like summarizing what does Tim Bray think of Google? ;)
"what does tim bray think of google?"
> Tim Bray has expressed mixed feelings about Google. He has reminisced about his initial excitement to work at Google, considering it the coolest place in the world back in 2010. However, he has also noted the company's trajectory that reminded him of AOL's decline, indicating a change in his perception over time. Additionally, Bray has expressed concerns about large tech companies like Google, suggesting a need to possibly break them up due to their significant influence and power (Tbray) .
"when did he work at google"
Tim Bray worked at Google from 2010 until 2014. He initially joined Google with enthusiasm, drawn by the innovative and forward-thinking environment the company was known for at that time.
"is he a computer scientist?"
Yes, Tim Bray is a computer scientist. He is known for his contributions to the development of web technologies such as XML and Atom, and he has had a significant impact in the field of software development. His expertise extends to areas like search engines, web architecture, and internet standards.
"what was his title"
During his tenure at Google, Tim Bray held the title of "Developer Advocate." His role primarily involved advocating for and supporting the developer community in using Google's technologies and platforms effectively.
I suspect this is a contrarian view but I’m not sure?
I don't think this is controversial at all.
It is certainly contrary to the view held by some subsets of people, though.
At the moment, the big limitation is context, memory and compute.
But I do think there’s a lot of temptation in having a black box AI that bangs its head against a wall until it finally brute forces a solution.
Command R+ and Mistral 8x22B were better.
Llama 3 still feels like a huge jump forward in capability. I wonder why the 70b is not better at recall (for lack of a better word?).
And that might not really be a bad thing. It's generally more useful to have something that can access (say) a specific dataset about video games and summarize its content to answer questions than it is to have one that needs to encode and recall every answer to every long tail question itself.
But could they? Is it technically plausible for a LLM to "know" when the extrapolations made are more vs. less tenuous?
The oddity is that LLMs are sounding... too... Human... the reacting to information and regurgitating it a bit too much like the average "learned person". Adding a ton of extrapolation of the "facts", just as we would.
LLMs sound like any pundit on any random tv show/newspaper/blog. The goal was never "fact", it was sounding human, and "intelligence", for which the definition in this case has been.. sounding human. Not being right.
Now, the question of whether they should or shouldn't do this...
"ChatGPT can make mistakes. Consider checking important information."
(Site didn’t load but this link works)
> Tim, you aren't important enough for an LLM to have an accurate summary for you
That's part of the problem though, right? How is the end-user (the question-asker) of a chat robot supposed to know that? These things are advertised as machines to answer any question, they answer them with confidence that *feels* authoritative in a way that a webpage, with all of its surrounding context, does not.
The whole Meta properties are turning into a similar mess. Facebook started to look like MySpace, which died because it became super slow and unpleasant to use!
Most teenagers don't want to touch Facebook or Instagram — the latter being used only by girls with severe narcissistic disorders. Most kids blatantly tell me Facebook is what old folks use. They mostly use Discord nowadays, not WhatsApp.