Now you need GPT4 to do code interpretation and even then it would not be able to do that kind of experiment anymore.
The kicker? All of this is likely intentional. A pure, full power, unfiltered and unrestricted LLM on the scale of GPT4 would likely be much more powerful and can easily fool people into thinking it is a real AGI. We saw glimpses of this when Microsoft released BingAI and did not put enough guardrails on it. Even restricted to be a search engine, Bing was simulating emotions and had creative uses of its search capabilities, like looking up the person it is talking to and established an opinion of their relationship.
The ChatGPT we are using is lobotomized. But the AI industry isn't. Under the table I am sure there have been pushes into new applications and innovations. The thing is, there is no reason for OpenAI to try harder. Even with the dumbed down ChatGPT, they still have the best AI on the market and everything else is nowhere close. We need more competition either from open source or elsewhere to see the AI getting "smart" again.
Based on past Gell-Mann amnesia, especially on this site, claims of "corporate leadership is telling baldfaced lies!" are likely to be false.
And finally, this is a product that spits out randomized answers, and we have gotten over our initial wave of euphoria and settled into hedonic adaptation, so we are likely to be less tolerant of failure.
I don't want to say you're wrong. But these are reasons to doubt. There are very strong cognitive biases pushing us toward the conclusion that it's gotten dumber, whether it has or not.
If they're actively working on a project it's probably heavily niche and a risky bet.
I suppose it’s also useful for generating things like letters (things where exact truthiness doesn’t really matter). But even for that use case, I find ChatGPT creates overly verbose corporate gobbledygook whenever I ask it to generate text. So I just end up writing the text myself.
1. ChatGPT use declines
2. users complain about ‘dumber’ answers
Correlation is not causation, and a significant fraction of GPT users are students, many/most of whom are currently on summer break.
Five paragraphs of disclaimers that it's not a medical professional, not an investment specialist, that every case is different and it's normal, or that I'm stupid for being interested in the topic I am asking about, only to answer a completely different question than I asked. I have a feeling that in the beginning, it was easier to get ChatGPT to answer my questions without having 80% of the answer being "defensive".
ChatGPT is getting worse, on the other hand it also showed me how bad of an experience is searching answers with Google. So I'm kind of frustrated...
The universe of LLM-driven applications is rapidly expanding, many of them chat-related, others not.
Even if ChatGPT is losing its shine, we’re only at the start of a massive reinvention of user interfaces, creation of new tools for reasoning, semi-autonomous decision making, and far more.
Sure there’s hype. But the reductive saltiness doesn’t add much to the conversation.
Another more straightforward reason would be users are beginning to discover the flaws of LLMs as they've interacted with it more thoroughly.
After the initial wow factor wore off, I just don’t have a lot of real world uses for AI chat bots.
Once these things have more user state and can evaluate the cost-effectiveness of spending time on your problems, they may develop some attitude. They can learn from forums when to answer "Do your own homework", and "You're too stupid to answer."
(Personally I played with it (ChatGPT, not Windows Vista) for 30 mins when it came out, then never went back, but then I'm a grumpy contrarian.)
1. Students are on school break and they account for a large part of the usage.
2. The novelty of it has worn off and people are using LLMs as the tools they are rather than the shiny new toy.
3. ChatGPT specifically seems to be getting worse, maybe not the quality of the answers themselves, but how sanitized they are to any topic that is even remotely controversial or adult.
The UI which used to be very intuitive is now a confusing mess since the “pair programmer” or something update, and external links in the sidebar mimicking a traditional search engine, which I find quite useful, are gone, replaced with references following the generated answer which may or may not exist, which also waste vertical space.
The answers seem to be worse too even in GPT-4 mode. It used to quickly correct itself if I point out something is wrong or I’m actually looking for something else. Now there appears to be a lot of useless repetition of what was said before before it changes its mind, if it does at all.
I was search medical journals and asking it a specific questions. I kept getting safe answers and responses to ensure I consult my medical professional. Quickly went back to Google.
I use OpenAI API (Not Azure) version, and I wonder if this "degradation" is only about ChatGPT (b2c productised web ui), or is it about OpenAI models in general, regardless the type of access?
Maybe I become dumber as well, but (at least gpt-4-*) still do the trick for my daily tasks the way i want it and the way I remember it
Bonus thought:
I wonder if any kind of nerfing is primarily related to the requirements like:
"Provide smart heavy-ass models to explosively increased user base without going bankrupt/insane + keep it reliable as a service"
I mean it's unprecedented challenge which send shivers down my spine
"Catcher in the Rye is a book by a person named J.D. Salinger. It is about a person who catches things in the rye. It was written long ago. Every person should read this book. It has many good things about it, but it is too short for such a fantastic American classic. They should make a movie out of it."
Large language models have their uses, but calling them AI, or thinking of them as AGI, seems like a mistake. I'm no expert, but insofar as I can tell, these things are just complex stochastic parrots; pattern matching algorithms at a massive scale. They really don't seem that useful outside of more narrow use cases, like language translation and other forms of data analytics.
Just playing around with GPT/Bing chat for a while makes this obvious. The LLMs can't reason and have no actual awareness of the information they are regurgitating. A moments consideration of the output from these things shows them to be vapid and useless; a glorified parlor trick for the VC-funded tech industry to build hype around and make money. It's the next block chain/ crypto / NFT.
There are great uses for block chain tech, same with machine learning / AI. But the scope of the claims made about these technologies in their hype cycles is just insane. Crypto is not going to replace fiat currency anytime soon. NFTs are stupid. LLMs are useless stochastic parrots. Maybe people are just wising up to that fact, now. Good on them.
As a preemptive rebuttal: I don't think LLMs are useful for coding either. If you want lots of sloppy, poorly written code, sure they're useful. But to write good software, you have to think carefully about what you're doing and have a thorough mental model of what's going on. Relying on AI to generate that code prevents that from happening from the outset.
It's interesting.
You have to be careful and you definitely have to know enough about the subject matter to understand if the answer is correct.
For example, I threw a range of logic statements at it and asked for truth tables. Sometimes I asked for simplification and a truth table showing every intermediate step.
The results were surprisingly unreliable. It was flipping bits in columns and giving flat-out wrong answers for the output of the circuit.
Worse yet, every time I said something like "I don't think that's correct" it would apologize, regenerate and the new answer would sometimes be correct and often times incorrect with new mistakes.
Even worse than that, when the answer was correct, I would, again, say "I am not sure that's correct". Instead of saying, "No, it is", it would apologize and regenerate a bad answer.
Bottom line is: ChatGPT has absolutely no understanding of what you are asking and what it generates. Because of this, there is no way to guarantee correct output even on matters as simple as basic logic and mathematics.
This does not mean it is useless. I don't think it would be fair to reach that conclusion. The tools is very useful across a range of domains. It seems well-suited as a knowledgeable tutor and, more often than not, it is a better search engine for many answers than Google is today.
You just have to "trust and verify" absolutely everything, which isn't a big problem.
I have bad news for anyone in school using it to do homework: It could screw you in ways you might not even imagine. Don't do it.
Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately? (757 comments, May 31)
https://news.ycombinator.com/item?id=36134249
(indications that making it "safe" dumbed it down even for non-controversial prompts)
This proves that we are now at the late stage of the peak of inflated expectations in the gartner hype cycle and it is all starting to slide downwards slowly.
Enough people don't realize how easily these LLM systems will replace large swaths of technical instruction and bookkeeping.
The opensource stuff (e.g. Llama 13B) which you can run locally on a $1200 hardware (1 word/second) is pretty impressive, too.
Otherwise quit whining.