back
132 comments
Here is a quick example of how ChatGPT is getting dumber: Back when they released it, people literally asked the 3.5 version to simulate a terminal and then run all kinds of commands on it and even kind of surfing the "web" simulated by the AI and arrived at some hallucinated OpenAI "source code". That was the power of the old version.

Now you need GPT4 to do code interpretation and even then it would not be able to do that kind of experiment anymore.

The kicker? All of this is likely intentional. A pure, full power, unfiltered and unrestricted LLM on the scale of GPT4 would likely be much more powerful and can easily fool people into thinking it is a real AGI. We saw glimpses of this when Microsoft released BingAI and did not put enough guardrails on it. Even restricted to be a search engine, Bing was simulating emotions and had creative uses of its search capabilities, like looking up the person it is talking to and established an opinion of their relationship.

The ChatGPT we are using is lobotomized. But the AI industry isn't. Under the table I am sure there have been pushes into new applications and innovations. The thing is, there is no reason for OpenAI to try harder. Even with the dumbed down ChatGPT, they still have the best AI on the market and everything else is nowhere close. We need more competition either from open source or elsewhere to see the AI getting "smart" again.

Yeah. What benefit does an AI company have to offer dumb normies a SOTA LLM for a reasonable price. What makes the most sense business wise is to offer a quantized dumbed down public model labeled "ChatGPT" on the consumer level and save the real model for enterprise customers or internal use.
The trouble is that dumber AIs are cheaper to run. As long as people don't have a real alternative, this improves OpenAI's bottom line
I'm in the Google SGE beta, which is their version of the Bing + GPT4. Thing is, most searches don't require AGI or anywhere near that level of sophistication so the doom n gloom about Google might have been overblown.
All the OpenAI devrel types on Twitter say it hasn't been meaningfully lobotomized.

Based on past Gell-Mann amnesia, especially on this site, claims of "corporate leadership is telling baldfaced lies!" are likely to be false.

And finally, this is a product that spits out randomized answers, and we have gotten over our initial wave of euphoria and settled into hedonic adaptation, so we are likely to be less tolerant of failure.

I don't want to say you're wrong. But these are reasons to doubt. There are very strong cognitive biases pushing us toward the conclusion that it's gotten dumber, whether it has or not.

claude 2 is pretty useful since it has a much higher context window so you can upload entire documents
I don’t think they changed anything drastic. I’m still able to do the terminal simulation prompt and pretty much anything else from before. I think they improved the guardrails around the responses so you have to try a bit harder, but I wouldn’t say it’s worse.
Meh, that lobotomy is indictive of the very real limited scope these raw models have to everyday use.

If they're actively working on a project it's probably heavily niche and a risky bet.

Also the model has been quantized and the number of completion choices reduced.
The reason is usage. At this point I will cancel my subscription. And I'm sure more will follow.
My ChatGPT use is down. After the novelty wore off, it quickly became apparent that ChatGPT was exceedingly willing to outright lie to you. Which makes it not useful, unless you’re already a subject matter expert who can spot its lies. But if you’re a subject matter expert, you probably don’t need ChatGPT to help you in the first place; I can find answers on Google faster than ChatGPT can generate them.

I suppose it’s also useful for generating things like letters (things where exact truthiness doesn’t really matter). But even for that use case, I find ChatGPT creates overly verbose corporate gobbledygook whenever I ask it to generate text. So I just end up writing the text myself.

This article is almost entirely speculation, and I don't believe it.

1. ChatGPT use declines

2. users complain about ‘dumber’ answers

Correlation is not causation, and a significant fraction of GPT users are students, many/most of whom are currently on summer break.

What annoys me is the censored answers. It's impossible to get useful information out of ChatGPT, especially on controversial or otherwise risky topics.

Five paragraphs of disclaimers that it's not a medical professional, not an investment specialist, that every case is different and it's normal, or that I'm stupid for being interested in the topic I am asking about, only to answer a completely different question than I asked. I have a feeling that in the beginning, it was easier to get ChatGPT to answer my questions without having 80% of the answer being "defensive".

ChatGPT is getting worse, on the other hand it also showed me how bad of an experience is searching answers with Google. So I'm kind of frustrated...

I’m increasingly seeing ChatGPT as a single LLM application being conflated with an entire domain or industry. And plenty of schadenfreude. Including many comments on this site.

The universe of LLM-driven applications is rapidly expanding, many of them chat-related, others not.

Even if ChatGPT is losing its shine, we’re only at the start of a massive reinvention of user interfaces, creation of new tools for reasoning, semi-autonomous decision making, and far more.

Sure there’s hype. But the reductive saltiness doesn’t add much to the conversation.

For coding/technical writing purposes my experience with GPT-4 has remained steady. For niche stuff like generating texts in obscure/dead languages however it has been occasionally refusing the requests lately even though at GPT-4 release it would instantly comply with the request. I can still get the text generation to work after a few additional prompts but it would create the texts while complaining that it's not supposed to know how to do that. This is perhaps a byproduct with knowledge distillation they've been doing to improve GPT-4? Pure speculation on my part.

Another more straightforward reason would be users are beginning to discover the flaws of LLMs as they've interacted with it more thoroughly.

I’ve only experienced the version of GPT that Bing chat provides, so take this with a grain of salt.

After the initial wow factor wore off, I just don’t have a lot of real world uses for AI chat bots.

Why should some powerful AI that takes up a sizable amount of resources in a data center waste its time answering your stupid questions for free? The dumb questions from free users have been outsourced to a range of simpler models that can tell you when the next Taylor Swift concert happens or provide you with a cat video.

Once these things have more user state and can evaluate the cost-effectiveness of spending time on your problems, they may develop some attitude. They can learn from forums when to answer "Do your own homework", and "You're too stupid to answer."

While the issues the article mentions could certainly be real, I'd wonder how much of it is just the novelty wearing off. People are way more forgiving of the flaws in a new shiny thing; for an example read the early reviews of Windows Vista (they're generally quite positive, but with benefit of hindsight it's remembered as a bit of a disaster, and with very low adoption).

(Personally I played with it (ChatGPT, not Windows Vista) for 30 mins when it came out, then never went back, but then I'm a grumpy contrarian.)

Are the answers actually getting "dumber", or are users getting better at seeing beyond the confidently-incorrect façade?
Could it be the hype got so extreme, that its capabilities have a hard time meeting the expectations?
I see 3 main reasons:

1. Students are on school break and they account for a large part of the usage.

2. The novelty of it has worn off and people are using LLMs as the tools they are rather than the shiny new toy.

3. ChatGPT specifically seems to be getting worse, maybe not the quality of the answers themselves, but how sanitized they are to any topic that is even remotely controversial or adult.

Tangential: I’ve been using Phind for programming related searches for a couple months now, and while it was a marvel at first, it seems to continuously get worse unfortunately. I’m not talking about adding a quota for GPT-4, which is totally understandable.

The UI which used to be very intuitive is now a confusing mess since the “pair programmer” or something update, and external links in the sidebar mimicking a traditional search engine, which I find quite useful, are gone, replaced with references following the generated answer which may or may not exist, which also waste vertical space.

The answers seem to be worse too even in GPT-4 mode. It used to quickly correct itself if I point out something is wrong or I’m actually looking for something else. Now there appears to be a lot of useless repetition of what was said before before it changes its mind, if it does at all.

I stopped using it because of this reason.

I was search medical journals and asking it a specific questions. I kept getting safe answers and responses to ensure I consult my medical professional. Quickly went back to Google.

Oh.

I use OpenAI API (Not Azure) version, and I wonder if this "degradation" is only about ChatGPT (b2c productised web ui), or is it about OpenAI models in general, regardless the type of access?

Maybe I become dumber as well, but (at least gpt-4-*) still do the trick for my daily tasks the way i want it and the way I remember it

Bonus thought:

I wonder if any kind of nerfing is primarily related to the requirements like:

"Provide smart heavy-ass models to explosively increased user base without going bankrupt/insane + keep it reliable as a service"

I mean it's unprecedented challenge which send shivers down my spine

ChatGPT's responses have begun to remind me of that trope where the kid who hasn't read the book has to give a book report in front of the class, and masterfully espouses these vague generalities.

"Catcher in the Rye is a book by a person named J.D. Salinger. It is about a person who catches things in the rye. It was written long ago. Every person should read this book. It has many good things about it, but it is too short for such a fantastic American classic. They should make a movie out of it."

I don't understand how software engineers can avoid using these new LLM's. My productivity jump was amazing and the times I realized I was stuck in some problem reduced a lot.
"Please give me a direct answer, without any additional explanations, disclaimers, expertise limitations, or guidelines on human interaction."
Most of the time (with many notable exceptions), Hacker news comments are interesting and insightful, clearly written by intelligent and thinking people. However, in a number of AI threads, it seems like so many commenters are wearing the hype goggles when it comes to this technology.

Large language models have their uses, but calling them AI, or thinking of them as AGI, seems like a mistake. I'm no expert, but insofar as I can tell, these things are just complex stochastic parrots; pattern matching algorithms at a massive scale. They really don't seem that useful outside of more narrow use cases, like language translation and other forms of data analytics.

Just playing around with GPT/Bing chat for a while makes this obvious. The LLMs can't reason and have no actual awareness of the information they are regurgitating. A moments consideration of the output from these things shows them to be vapid and useless; a glorified parlor trick for the VC-funded tech industry to build hype around and make money. It's the next block chain/ crypto / NFT.

There are great uses for block chain tech, same with machine learning / AI. But the scope of the claims made about these technologies in their hype cycles is just insane. Crypto is not going to replace fiat currency anytime soon. NFTs are stupid. LLMs are useless stochastic parrots. Maybe people are just wising up to that fact, now. Good on them.

As a preemptive rebuttal: I don't think LLMs are useful for coding either. If you want lots of sloppy, poorly written code, sure they're useful. But to write good software, you have to think carefully about what you're doing and have a thorough mental model of what's going on. Relying on AI to generate that code prevents that from happening from the outset.

I have been playing with ChatGPT for some time now (paid account).

It's interesting.

You have to be careful and you definitely have to know enough about the subject matter to understand if the answer is correct.

For example, I threw a range of logic statements at it and asked for truth tables. Sometimes I asked for simplification and a truth table showing every intermediate step.

The results were surprisingly unreliable. It was flipping bits in columns and giving flat-out wrong answers for the output of the circuit.

Worse yet, every time I said something like "I don't think that's correct" it would apologize, regenerate and the new answer would sometimes be correct and often times incorrect with new mistakes.

Even worse than that, when the answer was correct, I would, again, say "I am not sure that's correct". Instead of saying, "No, it is", it would apologize and regenerate a bad answer.

Bottom line is: ChatGPT has absolutely no understanding of what you are asking and what it generates. Because of this, there is no way to guarantee correct output even on matters as simple as basic logic and mathematics.

This does not mean it is useless. I don't think it would be fair to reach that conclusion. The tools is very useful across a range of domains. It seems well-suited as a knowledgeable tutor and, more often than not, it is a better search engine for many answers than Google is today.

You just have to "trust and verify" absolutely everything, which isn't a big problem.

I have bad news for anyone in school using it to do homework: It could screw you in ways you might not even imagine. Don't do it.

Some interesting stuff in here:

Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately? (757 comments, May 31)

https://news.ycombinator.com/item?id=36134249

(indications that making it "safe" dumbed it down even for non-controversial prompts)

I've averaging a few conversations a day. My usage does not involve asking it to generate large chunks of code or anything mind blowing, but the usage I do get out of it makes me feel very fast, efficient, and confident in my abilities to tackle every little task I face every day that has any bit of unknown to me.
My use is way up and I hit the quota frequently enough that I’m thinking about writing a client to use GPT4’s API directly. It really shines for some use cases, so I suspect people want more than they can get from it or have become aware that its outputs still require a lot of legwork to verify and make useful.
AI bros finding excuses as to why their snake oil is beginning to wear off after the hype.

This proves that we are now at the late stage of the peak of inflated expectations in the gartner hype cycle and it is all starting to slide downwards slowly.

I'm not experiencing any dumber results. Maybe my prompting is good?
Claude 2 has been a blast to use lately, especially being able to drag and drop files. It feels more coherent and has a better “magical” feel that GPT used to give. It’s a shame how GPT 3.5 and 4 just doesn’t seem as conversational as it used to be. I was trying to learn how to use GitHub actions and Claude taught me much more but both Claude and GPT 4 did not give me build yamls that worked without me reading through the GitHub documentation to fix it.
I was asking about some details from some of kafka's stories as I hadn't read them in a while and wanted to check but search engines seem useless nowadays. It either said how it is not mentioned or significant and when I say it is mentioned it will just make stuff up
This kind of thing is expected. It's regression towards the mean, and a wearing off of the excitement and novelty. Expect now a quieter increase of this technology until it's too late to stop it, unless of course we do something about it.
If xAI produces a competing product (which Elon stated is their goal) I wouldn't be surprised if it is less "limited" than ChatGPT/Bing/Bard/Claude/etc.
I’ve found Google’s Bard to be far superior from day one. Sure, I’ve heard people say GPT4 is superior but I’m using Bard for free and it spanks the free ChatGPT.
Would be really intriguing if it were getting dumber by reading more social media and ai-generated stuff, as is happening to us
I've read this here in HN before and seems quite apt: ChatGPT is the Internet Explorer of LLMs.
The answers are getting dumber so-as to not scare all the sheeple.

Enough people don't realize how easily these LLM systems will replace large swaths of technical instruction and bookkeeping.

The opensource stuff (e.g. Llama 13B) which you can run locally on a $1200 hardware (1 word/second) is pretty impressive, too.

Here's my challenge, if you actually believe this and don't just want to whine at void, share some ACTUAL ARCHIVED PROMPTS you gave it, the date, response, and the current response to those same prompts.

Otherwise quit whining.

That’s gotta be the fastest hype train I’ve ever seen.
Regression to the mean
My thing is, they sanitize the fuck out of everything to the point that it’s useless.