back
59 comments
I wonder if there is any connection between the models producing exaggerated outputs and the litany of exaggerated or overconfident claims that academic media offices or the press have produced from previous studies. Maybe the models trained on the studies and the reports on the studies naturally tend toward the style of attention seeking reports even when directly provided with the studies.
This is the same mistake we were seeing in commercial use of AI.

   1. "This process is flawed due to human bias"
   2. Train AI/ML to make the same decisions with the same outcome
   3. "How can there be any flaws in this process? AI is bias-free."
A repost of a previous comment, but the anthropomophization of this tech is so off the charts, I feel like I'm going to be repeating it, a lot:

One of the most offensive words in the anthropomophization of LLMs is: hallucinate.

It's not only an anthropomorphism, it's also a euphemism.

A correct interpretation of the word would imply that the LLM has some fantastical vision that it mistakes for reality. What utter bullsh1t.

Let's just use the correct word for this type of output: wrong.

When the LLM generates a sequence of words, that may or may not be grammatically correct, but infers a state or conclusion that is not factually correct; lets state what actually happened: the LLM generated text was WRONG.

It didn't take a trip down Alice's rabbit hole, it just put words together into a stream that inferred a piece of information that was incorrect, it was just WRONG.

The euphemistic aspect of using this word is a greater offense than the anthropomorphism, because it's painting some cutesy picture of what happened, instead of accurately acknowledging that the s/w generated an incorrect result. It's covering up for the inherent short comings of the tech.

Imagine the irony if this article were to exaggerate the claims made in the study itself.

Perhaps saying things like "Most leading chatbots routinely exaggerate science findings" instead of "We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet [...] with DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralizing in 26–73% of cases".

To be fair, the article itself already mentions this: "Summaries by models (1), (4), (8), and (9) didn’t significantly differ in the kind of generalisations they contained from the original text. “So basically,” Peters concludes, “Claude, in different versions, did really well.”"

I’m sure they were trained on clickbait articles and PR releases from universities, which routinely do the same and misinterpret and overstate the importance of the science.
I am shocked, I tell you I’m flabbergasted that LLMs which cannot truly reason go down the wrong rabbit-hole nearly every time. I’ve spent a good chunk of the weekend trying to accelerate the development of a small SaaS solution using Cursor, CoPilot, etc even with the latest and greatest Claude Sonnet 4, paying for “Max”, etc and any meaningfully sized request, winds up being a completely frustrating experience of trying to get these tools to stay on the rails. I’m to the point where I will give these fuckers explicit instructions to come up with several hypotheses, potential solutions, and then not to generate code before getting my go-ahead and it still more often than not go ahead and start editing code without my approval/direction. As a bonus it will forget a few things we corrected earlier. Can’t wait for the first vibe-coder to get sued when there’s a massive breach or financial loss. These tools are not ready for prime time and they certainly aren’t worthy of the billions being spent on them.
Absolutely, they're being hyped to the extreme talking about how it's AGI in the next two years, yet I cannot get one of them to generate a jar with no lid. It just creates the lid every single time, but that's just an example. They succeed at many things, but they fail so hard at such simple things that it casts a huge shadow on the times it does succeed.

Because they work with statistics and averaging training data, it makes sense that certain things it will have copied correctly, and others it will have completely mixed up and it cannot process properly. The problem is that they are literally being sold as reasoning models, and AI that can be just like a junior developer. These comparisons are borderline irresponsible as it could be driving a bubble that will burst and affect millions of people.

I'm sick of people saying LLMs don't reason as if they know unequivocally.

We dont' understand what's going on with LLMs. We have evidence of them reasoning successfully. And we have evidence of them Failing to reason.

That evidence DOES not logically lead to "LLMs don't reason". There are multiple possibilities here.

1. LLMs can reason, but they can't tell the difference between hallucination and reasoning.

2. LLMs can reason, but they choose to lie.

3. LLMs can't reason, when they get something right, it's pure coincidence.

There is not a single person who can prove or disprove ANY of those 3 points. What most people end up doing is ironically Identical to what the LLM does. They hallucinate an answer: "LLMs cannot truly reason."

Think about it. There is NO EVIDENCE or insight into how an LLM works that can even tell us how an LLM arrived at a specific response. We HAVE NOTHING.

Yet why do I see everywhere people like this guy, who makes claims out of nowhere? Which brings us back full circle: do humans reason? How similar is human hallucination to LLM hallucination?

Anti-AI people are kind of weird.

I got put on probation on SomethingAwful because I created a thread talking about an AI project I was working on (An Icecast radio station that uses OpenAI to generate DJ chatter and commercials), and everyone acted like I was simultaneously a completely uncreative moron also also somehow stealing work from people I would have hired, like I was somehow depriving a DJ of a job by not hiring one. I am an unemployed software person who is building something for fun, "hiring a dedicated 24 hour DJ" was never on the table.

In this thread, I noticed a lot of assertions that were just being accepted as axiomatically true, like asserting the AI is worse for the environment than humans doing the equivalent labor (which is not nearly as cut and dry), and that AI can't reason and that anything that involves any AI is inherently "theft". These assertions are completely unqualified and people just eat it up.

I don't know if LLMs "reason" by any consistent definition of the word. They might, they might not, I'm not going to pretend to know, but I find it a little irritating how people just assert that they don't and people just gobble it up.

op you asked, what about the energy expended to generate the art. the answer is this: the average human runs on about 2000-2500kcal of energy a day. if you were to spend 25 hours working on a piece of art, the energy expenditure is between 2000 and 2500kcal.

if you were to generate one piece of AI art, the energy expended is several magnitudes higher. but if you understood basic mathematics and human biology, you would have understood this going into the discussion

> Anti-AI people are kind of weird.

I'm not Anti-AI, I'm Anti-shit that doesn't provide meaningful value and for building reliable professional software...they are not as valuable as "they" would have you believe.

Yeah it’s annoying af. This is supposed to be a forum where more intelligent people gather but people just don’t get it.
I don't think the three options provided are exhaustive. For example:

4. LLMs can't reason, when they get something right, it's because the corpus of information used to create the model provides a very high probability that the response to that specific prompt is correct.

They're not exhaustive, but add as many options as you want to the ones given. Most of those options can't be proven or disproven. We don't understand what's going on with LLMs.

We've built something from scratch that we can't understand or control. That is what an LLM is.

Extraordinary claims require extraordinary evidence.

I don't have to prove that the LLM can't reason. The claims that they can have yet to be proven.

There was literally only One claim in this entire thread:

"LLMs can't reason"

No OTHER claim was made. So given the fact there was only one claim, Where does the burden of proof lie? Hint: the person who made the claim.

4. We don't actually have a coherent concept of what "reasoning" means in general. Same for "hallucination"...
We do. Reasoning is clearly defined. Sentience or consciousness is not.

Reasoning is basically the same as using logic to arrive at a conclusion when given a set of rules and axioms.

Hallucination is producing a conclusion by not following rules or axioms.

You can derive #3 from first principles of how the underlying gradient descent and backpropagation work. They approximate functions, therefore when it gets something wrong it means something went wrong with either the gradient descent or back propagation, or the training data.

We're not talking about getting a certain math question wrong (humans make mistakes like this). We're talking about ridiculous mistakes that are completely random and that humans would never make. I would even go as far to say that calling them mistakes is a stretch. The algorithm simply did not capture some mapping between the prompt and the output answer.

This is not how humans make mistakes. Humans build knowledge in some kind of logical knowledge tree, and mistakes arise from how this mechanism works (which we don't know exactly how, but we can certainly conclude it is completely different than transformers and LLMs). Most humans don't make random mistakes but rather something like logical mistakes.

tl;dr; It's clear to me LLMs make mistakes in such a way that exposes their simple underlying mechanism, as oppose to humans who make mistakes that contain rich layers of logical reasoning.

>You can derive #3 from first principles of how the underlying gradient descent and backpropagation work. They approximate functions, therefore when it gets something wrong it means something went wrong with either the gradient descent or back propagation, or the training data.

Then derive it from first principles. Use mathematical notation. You can't even draw this curve... it has so many dimensions. The crazy thing is you started hallucinating to me AFTER you told me you can derive it from first principles. You claimed you can derive it, then proceeded to NOT derive it.

>We're not talking about getting a certain math question wrong (humans make mistakes like this). We're talking about ridiculous mistakes that are completely random and that humans would never make. I would even go as far to say that calling them mistakes is a stretch. The algorithm simply did not capture some mapping between the prompt and the output answer.

Humans make plenty of ridiculous mistakes. Even so humans can lie. How do you know it's not lying? Again. Prove it.

>This is not how humans make mistakes. Humans build knowledge in some kind of logical knowledge tree, and mistakes arise from how this mechanism works (which we don't know exactly how, but we can certainly conclude it is completely different than transformers and LLMs). Most humans don't make random mistakes but rather something like logical mistakes.

Please derive how humans reason from first principles. I mean this by showing me experimental evidence that shows me the genesis of a signal traveling through the human brain and branching through billions of neurons to produce "reasoning".

Oh you can't? Well it looks like you're just making an approximation here? Possibly an Hallucination. Sound familiar?

>tl;dr; It's clear to me LLMs make mistakes in such a way that exposes their simple underlying mechanism, as oppose to humans who make mistakes that contain rich layers of logical reasoning.

You had to do make several assumptions and leaps in creativity to arrive at your conclusion. It's an hallucination through and through.

The is zero (legitimate) question about what LLMs do. They're math, not magic.

They highlight fun philosophical / definitional questions like the Chinese room thought experiment, but that's it.

And where is your evidence of this? Can you even prove this for a single query/response pair? Can you trace the genesis of the signals travelling through the neural network and prove to me via the structure that no reasoning was performed at all? You don't. You arrived at this claim with NO evidence.

How do you arrive at a claim without evidence? There's a word for it. It's called an Hallucination. Humans, like LLMs, have trouble saying "I don't know."

2 years of RLHF I don’t see reasoning. It’s clearly a text-similarity game, plus training of course. My mental model to work best is not that I’m interacting with a thing with reasoning. That differs from my mental model of interacting with humans. I would be worse at my job if I interacted with LLMs as if they reasoned.
Where's the evidence? You have none. Again you made a leap here with ZERO evidence which is equivalent to an hallucination.
RE: Claude Code, have you referenced this guide?

https://www.anthropic.com/engineering/claude-code-best-pract...

I have, and I still can't get it to stop acting like a drunken compulsive liar with an enormous wealth of information.
I've been wondering this as well. did you use any prompts to help split into roles of product, engineering, well defined tasks.
I even used Claude to help me formulate my prompts, Cursor rules, etc.
I think were just training them.

  “You asked her what color a house was and she said, ‘It’s white on this side.’”
  “That’s right.”
  “She didn’t assume that the other side was white, too… and a Fair Witness wouldn’t.” 
  -- Stranger in a Strange Land (1961)
An LLM is an abstraction machine, it mashes together anything that is nearby in a high dimensional space. Its statistical model is its source of truth. For a Fair Witness AI reasoning needs to supplant statistics. Which I'm guessing can get weird fast. LLMs are really good at being suggestible. For this we need the opposite.
That's an extremely interesting observation, it kind of reminds me of some of the traps bayesians sometimes fall into... (not saying bayesian reasoning can't be very useful)
"Over a year, we collected 4,900 summaries. When we analysed them, we found that six of ten models systematically exaggerated claims they found in the original texts"

So it turns out llms trained largely on Internet science articles make the same mistakes as are made by science journalists.

LLM being trained on a corpus of pitiful science reporting from the mainstream press... This is exactly what you would expect.
Studies like this should make it evident that LLMs are not reasoning at all. An AI that would reason like humans would also make mistakes like humans and by now we can all see that LLM mistakes are completely random and nonsensical.

It is also unclear that the current rate of progress is in the direction that would solve this issue. I think generative AI for images and video will get better, but the reasoning capabilities seem to be in a different domain.

> Studies like this should make it evident that LLMs are not reasoning at all. An AI that would reason like humans....

Humans don't reason either. Reasoning is something we do in writing, especially with mathematical and logical notation. Just about everything else that feels like reasoning is something much less.

This has been widely known at least since the stories where Socrates made everybody look like fools. But it's also what the psychological research shows. What people feel like they're doing when they're reasoning is very different with what they're actually doing.

Well no, most people can reason without writing or speaking. I can just think and reason about anything. Not sure what you mean.

Reasoning is something like structured thoughts. You have a series of thoughts that build on each other to produce some conclusion (also a thought). If we assume that the brain is a computer, then thoughts and reasoning are implemented on brain software with some kind of algorithm... and I think it's pretty obvious this algorithm is completely different than what happens in LLMs... to the extent that we can safely say it is not reasoning like the brain does.

There is also a semantic argument here, if we say that since we don't know what humans are doing then we can also stretch the word and use it for AI, but I think this is muddying the waters and creating all the hype that I think will not deliver what it's promising.

Most leading chatbots routinely exaggerate science findings - ahh yes this is completely unique to llms
To quote the article: "Chatbots were nearly five times more likely to produce broad generalisations than their human counterparts."

The following might be rude and unkind, but at this point frankly necessary: LLM apologists should stop projecting.

> Chatbots were nearly five times more likely to produce broad generalisations than their human counterparts

Never seen a journalist not make broad generalizations on a scientific study so I am not sure where these numbers are coming from.

What is the human counterpart to an LLM? The article doesn't describe.

And "LLM apologists" is so polemical it's hard to take seriously. We get it, you don't like GenAI. That's fine, but can we talk about it without getting normative?

Online forums for difficult medical conditions are full of ChatGPT copy-and-paste responses right now. LLMs are the tool of choice for people who want answers and want them now. Much like this article claims, a popular usage is to prompt the LLM to get a more optimistic interpretation about a specific supplement or treatment.

This is a really difficult problem because these people are often very sick and not getting the answers they want from their doctors. Previously this void was filled by alternative medicine doctors and quacks selling supplements. Now ChatGPT has arrived and has convinced a lot of them that they have a super-human AI at their fingertips that can be massaged to produce any answer they want to hear.

It’s painful to try to read some of these forums where threads have turned into endless “here’s what ChatGPT says” pasted walls of text, followed by someone else trying to counter with a different ChatGPT wall of text.

This isn’t unique to LLMs. There is a huge market for grossly exaggerating the conclusions of scientific studies. Podcasters like Huberman and Dr. Rhonda Patrick are famous for taking obscure studies with questionable conclusions and extrapolating to “protocols” or supplement stacks for their fans to follow. I often get downvoted when I mention fan-favorite podcasters by name, but I think by now many listeners have caught on to the way they exaggerate small studies into exciting listening material.

Definitely interesting how LLMs seem to be very good at reinforcing confirmation bias...
One of my friends went down a path of using ChatGPT to validate his belief that his alternative medicine approach was better than what his doctors recommended.

He had a whole host of tricks to work around ChatGPT’s protections or cautious replies. He’d strip out the cautions because he thought it was just OpenAI’s lawyers forcing disclaimers into the model. If he didn’t get the answer he wanted, he’d just retry or rephrase until he did.

I think the other half of the problem is that LLM users can be very good at pushing the LLM into doing confirmation bias. It’s much easier when you can hit the retry button with no side effects, unlike a human who will recognize what you’re trying to do when you keep asking different variations of a question until you get the answer you want.

Also, most leading news outlets routinely exaggerate science findings