It does sort of give me the vibe that the pure scaling maximalism really is dying off though. If the approach is on writing better routers, tooling, comboing specialized submodels on tasks, then it feels like there's a search for new ways to improve performance(and lower cost), suggesting the other established approaches weren't working. I could totally be wrong, but I feel like if just throwing more compute at the problem was working OpenAI probably wouldn't be spending much time on optimizing the user routing on currently existing strategies to get marginal improvements on average user interactions.
I've been pretty negative on the thesis of only needing more data/compute to achieve AGI with current techniques though, so perhaps I'm overly biased against it. If there's one thing that bothers me in general about the situation though, it's that it feels like we really have no clue what the actual status of these models is because of how closed off all the industry labs have become + the feeling of not being able to expect anything other than marketing language from the presentations. I suppose that's inevitable with the massive investments though. Maybe they've got some massive earthshattering model release coming out next, who knows.
They become extremely complex and sophisticated machines to squeeze a few more percent of efficiency compared to earlier models. Then diesel, and eventually electric locomotive arrived, much better and also much simpler than those late steam monsters.
I feel like that's where we are with LLM: extremely smart engineering to marginally improve quality, while increasing cost and complexity greatly. At some point we'll need a different approach if we want a world-shattering release.
The real typos were random missing letters. But the typos Gemini hallucinated were ones that are very common typos made in those words.
The only thing transformer based LLMs can ever do is _faking_ intelligence.
Which for many tasks is good enough. Even in my example above, the corrected text was flawless.
But for a whole category of tasks, LLMs without oversight will never be good enough because there simply is no real intelligence in them.
> It does sort of give me the vibe that the pure scaling maximalism really is dying off though
I think the big question is if/when investors will start giving money to those who have been predicting this (with evidence) and trying other avenues.Really though, why put all your eggs in one basket? That's what I've been confused about for awhile. Why fund yet another LLMs to AGI startup. Space is saturated with big players and has been for years. Even if LLMs could get there that doesn't mean something else won't get there faster and for less. It also seems you'd want a backup in order to avoid popping the bubble. Technology S-Curves and all that still apply to AI
Though I'm similarly biased, but so is everyone I know with a strong math and/or science background (I even mentioned it in my thesis more than a few times lol). Scaling is all you need just doesn't check out
In the meantime, figuring out how to train them to make less of their most common mistakes is a worthwhile effort.
If your expectations were any higher than that then, then it seems like you were caught up in hype. Doubling 2-3 times per year isn't leveling off my any means.
Compared to the GPT-4 release which was a little over 2 years ago (less than the gap between 3 and 4), it is. The only difference is we now have multiple organizations releasing state of the art models every few months. Even if models are improving at the same rate, those same big jumps after every handful of months was never realistic.
It's an incremental stable improvement over o3, which was released what? 4 months ago.
At this point it's pretty much given it's a game of inches moving forward.
Nothing in the current technology offers a path to AGI. These models are fixed after training completes.
In practical terms, Gpt 5 is a nice upgrade over most other models. We'll no doubt get lots of subjective reports how it was wrong or right or worse than some other model for some chats. But my personal (subjective) experience so far is that it just made it possible for me to use codex on more serious projects. It still gets plenty of things wrong. But that's more because of a lack of context than hallucination issues. Context fixes are a lot easier than model improvements. But last week I didn't bother and now I'm getting decent results.
I don't really care what version number they slap on things. That is indeed just marketing. And competition is quite fierce so I can understand why they are overselling what could have been just chat gpt 4.2 or whatever.
Also discussions about AGI tend to bore me as they seem to escalate into low quality philosophical debates with lots of amateurs rehashing ancient argument poorly. There aren't a hell of a lot new arguments that people come up with at this point.
IMHO we don't actually need an AGI to bootstrap the singularity. We just AIs to be good enough to come up with algorithmic optimizations, breakthroughs and improvements at a steady pace. We're getting quite close to that and I wouldn't be surprised to learn that OpenAI's people are already eating their own dogfood in liberal quantities. It's not necessary for AIs to be conscious in order to come up with the improvements that might eventually enable such a thing. I expect the singularity might be more of a phase than a moment. And if you know your boiling frog analogy, we might be smack down in the middle of that already and just not realize it.
Five years ago, it was all very theoretical. And now I'm waiting for codex to wrap up a few pull requests that would have distracted me for a week each five years ago. It's taking too long and I'm procrastinating my gained productivity away on HN. But what else is new ;-).
So yeah, maybe we are getting more incremental improvements. But that to me seems like a good thing, because more good things earlier. I will take that over world-shattering any day – but if we were to consider everything that has happened since the first release of gpt-4, I would argue the total amount is actually very much world-shattering.
The common concept for AGI seems to be much more about human replacement - the ability to complete "economically valuable tasks" better than humans can. I still don't understand what our human lives or economies would look like there.
What I personally wanted from GPT-5 is exactly what I got: models that do the same stuff that existing models do, but more reliably and "better".
Are you trying to say the curve is flattening? That advances are coming slower and slower?
As long as it doesn't suggest a dot com level recession I'm good.
I'd expect that at some level of reliability this could lead to a self-improvement cycle, similar to how a powerful enough model (the Claude 4 models in Claude Code) enables iteratively converging on a solution to a problem even if it can't one-shot it.
No idea if we're at that point yet, but it seems a natural use for a model with these characteristics.
> but given the types of things people have been saying GPT-5 would be for the last two years
This is why you listen to official announcements, not "people".This has me so confused, Claude 4 (Sonnet and Opus) hallucinates daily for me, on both simple and hard things. And this is for small isolated questions at that.
Is it actually simpler? For those who are currently using GPT 4.1, we're going from 3 options (4.1, 4.1 mini and 4.1 nano) to at least 8, if we don't consider gpt 5 regular - we now will have to choose between gpt 5 mini minimal, gpt 5 mini low, gpt 5 mini medium, gpt 5 mini high, gpt 5 nano minimal, gpt 5 nano low, gpt 5 nano medium and gpt 5 nano high.
And, while choosing between all these options, we'll always have to wonder: should I try adjusting the prompt that I'm using, or simply change the gpt 5 version or its reasoning level?
Would been interesting to see a comparison between low, medium and high reasoning_effort pelicans :)
When I've played around with GPT-OSS-120b recently, seems the difference in the final answer is huge, where "low" is essentially "no reasoning" and with "high" it can spend seemingly endless amount of tokens. I'm guessing the difference with GPT-5 will be similar?
This is sort of interesting to me. It strikes me that so far we've had more or less direct access to the underlying model (apart from the system prompt and guardrails), but I wonder if going forward there's going to be more and more infrastructure between us and the model.
No mention about the (missing) elephant on the room, where are the benchmarks?
@simonw has been compromised. Sad.
> -------------------------------
"reasoning": {"summary": "auto"} }'
Here’s the response from that API call.
https://gist.github.com/simonw/1d1013ba059af76461153722005a0...
Without that option the API will often provide a lengthy delay while the model burns through thinking tokens until you start getting back visible tokens for the final response.
> and minimizing sycophancy
Now we're talking about a good feature! Actually one of my biggest annoyances with Cursor (that mostly uses Sonnet).
"You're absolutely right!"
I mean not really Cursor, but ok. I'll be super excited if we can get rid of these sycophancy tokens.
right :-D