back
213 comments
Probably too late for this, but I have argued before that language is a fundamentally lossy encoding of the human experience. We do our best to describe what we're seeing and experiencing using language, which is fantastically expressive, but it has its limits. I think we see glimpses of this when we find ourselves saying things such as, "it's impossible to put it into words" or we overload certain words when we mean very different things, such as, "I love my children" or "I love apple pie". Clearly the word "love" here has a certain magnitude that is not being expressed, yet it is understood by the listener somehow.

So, I do sort of buy into this idea that Einstein was simulating the world and running experiments on those simulations in ways that were beyond what you could encode in natural language. Will AI be capable of doing this, if it is bounded by training data that is composed almost entirely on language? One might argue that if AI is training on a lossy encoding/representation of the human experience, how will it be able to simulate anything beyond that experience? Unless it does so in a way that we manage to do when we image objects beyond 3D. But now I'm just rambling.

The lossy compression of language is why we should find it unsurprising that LLMs tend to perform better at code, than at human language tasks or reasoning. While there can be subtle semantic differences in real codebases (using "null" to mean "unknown" in one context, versus "intentionally blank" in another), there is a much tighter coupling of semantics to meaning (low ambiguity) compared to "love" in English (let alone any inexpressible je ne sais quoi).

With apologies if this is common knowledge at this point, 3Blue1Brown has been doing an excellent series on compression, and its relationship to intelligence (or more controversially, that they are one and the same): https://www.youtube.com/watch?v=l6DKRf-fAAM

But that also throws in sharp relief, that there is vastly more to the human experience than intelligence alone: qualia, desire, gut instinct, intuition, emergent creativity. (Whether the "God of the Gaps" for the delta between capabilities of human vs AIs is fixed, or diminishing, or even shrinking to zero, remains an open experiment we're all living through.)

> Probably too late for this, but I have argued before that language is a fundamentally lossy encoding of the human experience. We do our best to describe what we're seeing and experiencing using language, which is fantastically expressive, but it has its limits.

Yes, and sometimes this is very intentional. Take for example a short poem which if you sit and really think about it for a long time, you could go off on a mental tangent of imagining what sort of kingdom or empire created a statue that is now "two vast and trunkless legs of stone", for instance. Being terse and allowing for human interpretation is kind of the entire point of something being written like this.

I met a traveller from an antique land

Who said: Two vast and trunkless legs of stone

Stand in the desert. Near them, on the sand,

Half sunk, a shattered visage lies, whose frown,

And wrinkled lip, and sneer of cold command,

Tell that its sculptor well those passions read

Which yet survive, stamped on these lifeless things,

The hand that mocked them and the heart that fed:

And on the pedestal these words appear:

"My name is Ozymandias, king of kings:

Look on my works, ye Mighty, and despair!"

Nothing beside remains. Round the decay

Of that colossal wreck, boundless and bare

The lone and level sands stretch far away.

In this context llm's remind me of the anecdote of Agassiz and the fish:

"A post-graduate student equipped with honours and diplomas went to Agassiz to receive the final and finishing touches. The great man offered him a small fish and told him to describe it. Post-Graduate Student: “That’s only a sun-fish” Agassiz: “I know that. Write a description of it.” After a few minutes the student returned with the description of the Ichthus Heliodiplodokus, or whatever term is used to conceal the common sunfish from vulgar knowledge, family of Heliichterinkus, etc., as found in textbooks of the subject. Agassiz again told the student to describe the fish. The student produced a four-page essay. Agassiz then told him to look at the fish. At the end of the three weeks the fish was in an advanced state of decomposition, but the student knew something about it."

sourced from: https://nabeelqu.co/understanding

Or, ask yourself the following question:

If you read every single book there is about The Grand Canyon, and watched every single video and/or documentary about The Grand Canyon, do you believe that you have fully experienced The Grand Canyon? Or do you just have to be there to fully experience it.

I dunno. Substitute in whatever you want for "The Grand Canyon". Maybe climbing Mount Everest or walking on the Moon. The point is that maybe the human experience is more vast than what is written about it.

> The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be “voluntarily” reproduced and combined… The above-mentioned elements are, in my case, of visual and some of muscular type. Conventional words or other signs have to be sought for laboriously only in a secondary stage, when the mentioned associative play is sufficiently established and can be reproduced at will.

—Albert Einstein

Quoted in Using Spaced Repetition Systems to See Through a Piece of Mathematics,

https://news.ycombinator.com/item?id=18895613

which describes the author's experience that if you approach a field obsessively enough, eventually you begin to understand it at a level deeper than language.

If an LLM is big enough, I imagine something similar is happening.

Noah Smith had an interesting idea related to this in a recent newsletter:

https://www.noahpinion.blog/p/what-will-more-intelligence-ac...

> Another way of saying this is that there may be laws of the universe that humans can’t understand but AI can. I call these “cloud laws” — causal regularities that can be exploited by technology, but which are too diffuse and complex for an individual human being to either intuit or communicate. Human language seems to obey cloud laws, so why not other phenomena too? Perhaps social sciences like economics, sociology, and political science obey similarly complex regularities, and AI can help us find them. Perhaps there are physical processes — plasma, or topological materials, or aerial turbulence, etc. — that obey cloud laws instead of chaos?

I don't think that the 'lossy' part is the problem. Instead, the problem is that language is only a subset of the human experience. So there are aspects that are not included when training models based on language. Making the models multi-modal helps to close that gap, but it still exists.

But ultimately this is just a matter of training data. I do not say that it is easy to obtain the required data, but it is not a fundamental problem LLMs can't overcome.

This might be of interest:

Building AGI Using Language Modelshttps://bmk.sh/2020/08/17/Building-AGI-Using-Language-Models...

I think this is very critical concept and existing gap which does not makes into AI conversations. Interestingly enough a lot of Si Fi movies have captured this where the AI starts to feel and have "thoughts". Have to give kudos to the authors for being so creative and imaginative.
The popular retelling of how Einstein created Special Relativity to "Resolve the contradictions of Michelson-Morly experiments" is very reductive to the history of the question. The epitome is the quote from the paper:

> From the two postulates, Einstein derived the Lorentz trans- formation ...

If Einstein derived them, who is "Lorentz"?

The groundwork for Special Relativity was the study of electrodynamics and symmetries of Maxwell equations. The Einsteins paper was literally called "On the Electrodynamics of Moving Bodies" and never cites Michelson and Morley.

Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently:

> A few reflections on my "LLMs Can’t Jump" paper:

> My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.

> First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.

> This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.

> Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.

> Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.

> Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!

[1] https://x.com/TZahavy/status/2082401499628376180

Came for: "A computer once beat me at chess, but it was no match for me at kick boxing."

TFA was actually about leaps of intuition, sadly.

One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.

This is literally an opinion of one dude which is not backed by any kind of quantitative evidence.

It's actually possible to answer this question rigorously:

1. Define a scientific result which qualifies as a "jump". They should be frequent enough that they happen every year - otherwise one might say humans can't jump either.

2. Identify all such "jumps" in articles published in 2026, and use LLM with 2025 knowledge cut-off to re-derive these results with minimal amount of information.

It really irks me that people boost these low-effort articles just because they confirm pre-conceived notion that LLMs are limited

The theory is that creative leaps in theoretical physics require a grounding in sensory experience, but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding. They do address this at the end, saying

"In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality."

But if such sense experience is possible in abstract domains via some high-dimensional topology, why could a sufficiently advanced LLM not develop an equivalent high-dimensional topology for domains like physics and use it to make creative leaps?

Why everybody is obsessed with replacing humans with LLMs when it seems like the most profitable use cases (like coding agents) rely on enhancing human capabilities?

Until LLMs have some 0% error humans will have to be in the loop (even if they only serve to take responsibility of the process).

The paper is from the 30th of April this year, openAi announced the counter example to the unit distance problem on the 20th of May. That is to say this paper seems to have aged not much but quite poorly.
I think this is more of a function of the harness and the environment than the LLM. I've seen some LLM interactions over complex environments like Godot and Unity that challenges the notion that there is no "jumping" going on at all.

An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress.

True understanding requires not knowing, and LLMs cannot "not know". LLMs have to come up with an answer, this is their nature. They are search engines. We do have a similar mechanism; one can notice it by reflecting. The mechanism is an opposite of true thinking, as it merely looks up what is already "known". We "jump" when we temporarily turn this mechanism off.

That said, here's an experiment conducted by some Soviet psychologist, I forgot the name. The man wanted to study intuition. So he invented an experiment that was supposed to trigger it in laboratory conditions. (Take a moment to marvel at that; how would you approach such a task?) He gave people a few puzzles. One was to place some sticks according to some rules. Yet another was to find a path in a maze. The secret was that the path in the maze was the same figure as the solution to the stick puzzle.

And he observed interesting results. People who solved the maze after the sticks found the path much faster than the control group. If a subject was asked to comment how he was solving the maze, at the start or halfway through, the speed dropped to typical. Subjects normally didn't notice the similarities.

So there is something to study here, although it is obviously a case of pattern matching, only subconscious. This is a jump of sorts, but not the one I mean. What I mean is a Zen jump.

Best comment on this from 6 months ago: https://news.ycombinator.com/item?id=46870575
Maybe LLMs can't, but another form of AI will. I hope nobody is interpreting this as "nothing will never be as good as us".

I see similar thinking in stories of how humanity got here. Religion has thousands of years adapting to this problem, every time we explain something, the goal post moves. Catholics today accept evolution (or least the church does), but it is the "jump" from monkeys to humans where God is the only explanation.

Just 5 years ago we didn't have a technology that knows more about everything than even most experts. We keep coming up with benchmark after benchmark and LLM/AI keeps destroying them. Now we've moved the benchmark to "the jump". Again, maybe it's LLMs or the way we currently do them that can't do this, but eventually something will.

The article seems like an interesting Gedankenexperiment. However, I think it overrotates on the GR analogy.

For example "..ARC captures the logical leap, it misses the manipulative component—the physical sensation and embodied simulation..." makes lots of assumptions on how such a discovery must occur, e.g. through "physical sensation and embodied simulation". Results matter, not the path there.

For example, quantization of energy, at the core of QM, wasn't discovered through "physical sensation and embodied simulation" at all. Planck simply found that if energy is quantized, then one obtained the observed black-body radiation spectrum. There was no "physical sensation and embodied simulation".

I had a related insight, but in the domain of humor [1].

LLMs are inherently probabilistic, and there's currently no mechanism for producing an orthogonal directional change in the path traced through a latent space which is also contextually relevant (landing on a punch line).

In other words, LLMs are fundamentally incapable of making intuitive/orthogonal leaps in context.

It might be possible to add this capability with a new architectural component like transformers, but specifically for making "left turns"/intuitive leaps.

[1] https://wnmurphy.com/llms-cant-do-humor/

If you think LLMs can't "jump," watch Terence Tao engage with ChatGPT to try to understand the intellectual leaps achieved by Claude Fable 5 in finding a counterexample to the Jacobian conjecture https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed... . After that you may revise your opinion.
I have been writing a 'paper' [1] on an adjacent topic for months now. At some point, I decided to make it an empirical paper vs position paper. I am still chasing the experiments (when I get some free time waiting for agentic loops)

For this paper specifically, after reading the abstract [2], I felt almost certain that the author would have used Judea Pearl's ladder of causation (https://web.cs.ucla.edu/~kaoru/3-layer-causal-hierarchy.pdf) but they did not. Would have probably been a better argument to make.

[1] paper in quotes because it may never get published (it is over 20 pages atm). the core argument is that lack of native adjacency resolution makes problems harder and sample inefficient, not impossible

[2] "Using Einstein’s formulation of General Relativity as a case study, we demonstrate that LLMs are structurally incapable of creating new foundational axioms, particularly when observational data is scarce. "

Also, the claim that 'LLMs are structurally incapable of creating new foundational axioms' is provably false depending on where you place 'fundamental'.

It's a very shaky position, and the empirical track record of "LLMs can't..." is in itself a reason to call it into doubt.

Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation.

The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't jump" out there, as if "abduction" is an established class of problem with known computational properties and requirements that the LLM architecture fails to satisfy. It's none of those things - and the paper makes the claim without backing it by anything but rhetoric attempts at persuasion.

The proposed solution is also dubious. The empirical track record of dedicated "world models" for reasoning and problem-solving is, frankly, downright abysmal. Even integrating multimodal data into LLMs has failed to yield general reasoning capability gains.

LeCun's misadventures in the field aside, the main frontier lab that pushes in favor of "improving reasoning via multimodal fusion" is GDM - and Gemini isn't exactly a paragon of frontier reasoning capabilities. It has strong multimodal capabilities, but lags behind both OpenAI and Anthropic in performance outside that - while Anthropic is the lab that always treated multimodal grounding as an afterthought, and still trades blows with OpenAI at the very edge of the performance frontier. Multimodal grounding seems to work great as a way to improve an AI's ability to deal with those specific modalities, but it falters outside that.

Now, it's not impossible that everyone who tried multimodal world models for reasoning is just doing it wrong, and there is an undiscovered recipe for multimodal grounding that results in a step change in AI capabilities. But the results we have so far suggest it to be unlikely.

I've said this for a long time. This inability of AIs to generate novel explanatory hypotheses is a big blocker for their ability to automate jobs like accountant, middle manager, and even the lowly cashier.
Maybe modulating temperature can help here: have the LLM come up with ideas at high temperature, and then critique them at low.

This is also tied to halucinations: it is something that humans do (for writing fiction, and for "jumps") - but what LLMs currently lack is intellectual honesty. Coming up with bullshit is fine (and in this context valuable) - the important bit is putting those ideas through some form of rigor, or just immediately turn around and admit to talking shit.

So I'd arge that hallucinations are what prevent LLMs from doing this in a useful way.

I found this paper really thought provoking, but I think the conclusion of “world models are the solution” leaves something to be desired. People are already equipping agentic systems with physical simulation tools and exploring action-conditioned world models. This is cool because you can change the rules of the simulation and observe what happens, but it doesn’t address the core question of what to change the rules to, or even what the goal should be in the first place.
>"While Generative AI has mastered Induction (statistical pattern matching) and is rapidly conquering Deduction (formal proof), we argue it lacks the mechanism for Abduction—the generation of novel explanatory hypotheses."

Interesting!

Induction Vs. Deduction Vs. Abduction!

(You know, if you like Logic, Philosophy, Law, or... just plain different ways to think/reason about something! :-) )

There's no reason why LLMs can't be hooked up to a physical world feedback loop though, in fact that's already being done with chemistry research https://www.youtube.com/watch?v=AYSR02tcwes
For those who may not be aware this is a clever title derived from a film title https://en.wikipedia.org/wiki/White_Men_Can%27t_Jump
I feel like you could just add some noise or randomness to the LLM and start approximating the leaps that the human mind uses to solve and understand unrelated things. Maybe that’s naive, it’s just coming from my organic computer in my skull.
We just put different things together and then we evaluate it.

In math its simple: does the verification say its okay.

If its mechanical: is any property better than what we have already.

etc.

Don't these datasets already exist in terms of the real-world samplings we have when training robots to do every day tasks?
Even myself, I really can't remember a time where I had this "jump". Is very subjective to feel this jump
The proof is in the pudding, so far there isn't a proven (E=mc^2 type) breakthrough LLM's had made yet.
What I find absolutely fascinating about this paper is that recently I made the leap that physical representation was a necessary ingredient for invention based on my own experience (lack of abundance of evidence) So I intuitively agree with the premise. It's kinda meta.
The early physics background is messy and incorrect. I didnt read the full position paper, but from its start: The Lorentz transformations were by Lorentz, well before Einstein’s paper on special relativity; the principle of relativity also existed before the Einstein paper. The math was all there, with steps taken by Maxwell, Voigt, Larmor, Lorentz, and Poincare. Einstein supplied a clean physical interpretation, making all inertial frames equivalent, making simultaneity frame dependent, and explaining length and time deformations without the need of the concept of ether. Skimming the end of the paper with the arguments about lack of abduction or inability to make the analogy without sensory experience, I see that this paper is unfounded speculation rather than solid/hard philosophical logic. As a position paper it is OK to appear, but i think it misses the point of how LLMs or other autoregressive learners of future states can build analogies and intuition that can help them formulate new theories of the world. Soon it will be more obvious to everyone, so I am not very worried about these writings.
This tracks for me as someone making keeps of intuition in little-explored areas.

I just don't see any of the LLM users around at all. Clearly some force is guiding them all away from thinking any of the "leap of faith" thoughts that I am thinking.

Neither can I, if I'm being honest.
Is this a reference to the movie "White man can't jump"? If so they are in a surprise because the movie says otherwise.
I have a weird thought experiment: If you give a GPT-2/3 level LLM tools to search the internet - any document, can it build bigger, better LLMs?

You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses.

Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data.

LLMs don't have legs yet. If you're smart but can't touch things you only keep being smart
Can talkie figure it out?
Interestingly, halucinations might be the way to achieve that.