Personally I like to remind people that these things are next-token predictors, but then emphasize how truly astonishing the results we can get out of sufficiently advanced next-token predictors are.
> The tokens it "predicts" aren't sampled from any naturally occurring distribution; the model's output is the result of an optimisation process that rewarded behaviour that was useful, and that's fundamentally different.
https://arxiv.org/abs/2504.13837"Surprisingly, we find that the current training setup does not elicit fundamentally new reasoning patterns. While RLVR-trained models outperform their base models at small k (e.g., k = 1), the base models achieve a higher pass@k score when k is large. Coverage and perplexity analyses show that the observed reasoning abilities originate from and are bounded by the base model. "
Let's stipulate that what pretraining does is train next token prediction over a gigantic corpus. You can then sample from this distribution repeatedly (cf the Large Language Monkeys paper) and count how often it passes some deterministic verifier.
What GRPO-style RLVR does is precisely this, but then reward the trajectories which passed the verifier. These distributions are _by construction_ within the accessible output space of the pretrained model; you're reweighting the distribution so that pass@k goes up, because that's (for applications like programming) very useful. RLVR is about making sampling more efficient; the only new information being added to the system is the presence of the verifier, and note that you only get a reward when the verifier passes, so there's essentially no mechanism for "teaching new facts" here.
> Coverage and perplexity analyses show that the observed reasoning abilities originate from and are bounded by the base model
On the face of it this seems unsurprising given the policy gradient term directly minimises this difference.
I don't have a good feel for how the output of an RLVR-trained model concretely differs from the base model. My guess would be there are a fairly small number of "forks" where the training creates a token flip that sends the model down a more useful path.
The fact that the straight paths between the forks resemble the base model would again be unsurprising since (a) those are exactly the right context to continue to elicit more output that's relevant to solving the problem (so not penalised by RLVR), and (b) preservation drops naturally out of the policy gradient term you add to limit catastrophic forgetting in the base model.
Low perplexity could be explained by the relative sparsity of the forks in the output stream, and/or by forks already having high entropy in the base model. That also aligns with the pass-at-high-k: yes it's doing more exploration without training but it's a bit of a monkeys-on-typewriters situation.
Lack of novelty is readily explained by the fact that you need some nonzero pass rate in the base model to actually get some useful training signal from RLVR. That's a limitation of contemporary RLVR techniques, not a limitation on post-training in general.
I think there's room in that forks-and-straights characterisation for the RLVR'd model to be doing something that looks a lot like computation, while having low perplexity vs the base model. I don't see anything in my admittedly incredibly shallow skim of the paper that refutes that.
You can get into RL as part of explaining why it's so unnervingly good at picking a next token.
I guess it was more the "predictor" part I had issue with. There's a tendency to reach for statistical or probabilistic terminology to describe things that aren't usefully understood in those terms. For example in the "Speed Always Wins" LLM technical survey (https://arxiv.org/pdf/2508.09834):
> The gate is a crucial component to bring sparsity in MoE models. For a batch of input token representations X ∈ RT×D, the gate function G determines the probabilities of dispatching token xi to each expert e
...which is nonsense: the gate simply, directly, selects the experts. There's nothing probabilistic about it.
Isn't G a learned probability?
If you put something through a softmax the output is (trivially) a valid PMF. Does that matter? You're not sampling from it.
I'm not understanding, can you explain this more? How does it become more than a next token predictor? Isn't the post-training simply altering the sampled distribution? And isn't that distribution naturally occurring? It's the distribution of "useful" next token?
It’d be prediction if it’s “predict what would come next in this text sampled from distribution X”.
But what’s it predicting if we’re looking for new useful outputs? It’s finding a distribution that’s useful, and generating tokens, but it’s not predicting what comes next in a known sequence.
It tries to learn the distribution of "useful" results either through verified rewards or human feedback. Then it encodes that in the network. When you run inference later, it samples or selects from that distribution.
Maybe it's a matter of interpretation. It's not predicting the next token based purely on the training corpus's distribution anymore, the RL process fine tunes that distribution so it predicts the next token that is closer to what was rewarded during RL. But as I see it, it's still predicting the next token, just from a reenforcement learned distribution instead of one found in a corpus of data.
It's harder to believe something is conscious or threatening to achieve word domination once you understand that it's a machine that statistically figures out which word should come next.
The problem is, these models challenge our definition of "consciousness." Or at least they point out how hopelessly-inadequate our thinking on the subject is. Some people really, really don't like having their personal definition of consciousness challenged.
The correct response to "So what, it's just a next-token predictor" isn't a long dissertation on RLHF, training architectures, scaling laws and whatever, but rather to turn around and respond, "Sure, and how is that different from what we do?"
I recommend this explanation: https://www.astralcodexten.com/p/next-token-predictor-is-an-...
(Note: I don't actually think the consciousness question is the most important one in the near term. Where I think this line of reasoning gets really dangerous is when people use it to assert that LLMs can't or won't engage in certain behaviors no matter much they advance; this doesn't have anything to do with consciousness.)
But LLMs are only operating on text and humans are only operating on <waves hands>
At the risk of sounding overly flippant, all world domination has been achieved by some person(s) figuring out which word should come next. Words quite literally = action when it comes to LLM’s with tools access
It's exhausting to even consider where to begin addressing the assertion that good leadership is just predicting the next word to say. Especially considering the corpus available to most great leaders in history was extremely small. To think Hannibal's military campaigns were just because he'd read like ten books in his life and could accurately forecast effective rhetoric is...indescribably divorced from reality.
At the risk of seeming like a jerk.
Not to mention people more worried about whether the AI is motivated to hurt us than what human motivations can do with something that can autocomplete its way through every possible attack vector of cryptographic systems most of use would prefer remain secure.
(tbf I think the "glorified autocomplete" still works surprisingly well for programming outcomes too. Autocomplete [and fuzzy search of reference material] actually is useful and often right and certainly can save time even when it's only suggesting the rest of the variable name. But you might not want to commit everything it suggests...)
What does "glorified autocomplete" say in an ontological sense exactly? Nah. It's just a lazy dismissal.
BTW, autoregressive pretraining (autocomplete) is a part of training.
Practically every other attempt at describing them leans too technical and unfamiliar (stochastic parrot) or too anthropomorphic (even describing them as "not like a human" gets people thinking in terms of humans, like how if I mention that your tongue is in your mouth all the time, using up almost all of the room, feeling your teeth and tasting itself, you're now uncomfortably aware of it and the numerous bumps on the surface).
You need to work from a reference that has both a shared understanding, and does not lead to problematic "if X has Y, and Z is like X, then Z has Y" seemingly-logical derived beliefs. "Spicy autocomplete" is a fairly safe starting point in both ways.
They sound fairly human, until you notice the patterns. They sound like they're thinking, until you pay attention.