back
637 comments
It’s endlessly fascinating to read the AI transcript of an expert who _really_ knows how to cut to the chase. It just shows how much you can potentially squeeze out of these models. I’m also surprised to see that even Terrence Tao seems to use it in a way that resembles, in progression, how I use llms in my area of expertise (emphasis on progression and usage patterns, not absolute skill, obv I don’t match that): short pointed questions that goes all in on the jargon and machinery of the field and steers the llm hard (eg no softballs). I’ve noticed that llms switch their tone and meet you basically more or less on your level.
It reinforces how to "learn AI" is to first master the problem domain.

I can use AI for coding after decades of coding. I can't use it for theoretical physics because I can't evaluate the responses.

  > I’m also surprised to see that even Terrence Tao seems to use it in a way that resembles, in progression, how I use llms in my area of expertise
I didn't understood anything about the thread, but reading Terrence's messages was weird because it looked exactly like the discussions I have with LLMs

I've mostly seen people trying to oneshot a result, while I'll quickly experienced that going through steps/discovery was more effective and more satisfying, since you can always steer it back in the right direction; while oneshotting is hit (and it kind feel like magic) or miss (and you'll have to rework your prompt).

I encountered something fairly similar working with Claude a few days ago. For a current project I've been fairly hand-wavy with requirements since I was getting good results, but it seemed to be failing hard on some key points, so I started to be more strict with it. Even after the fails were resolved, I've noticed that Claude now behaves differently within that project, carefully checking and rechecking things up front and also looking to me for guidance more often. Mildly irritating, but if it works...
The top comment has a counter example though https://x.com/DmitryRybin1/status/2079904005652893709

I think it's not even about the ability to steer the AI. Just the ability to ask the right questions

Came to say the same thing. Experts in any field have a huge unlock from AI.
Is there a site with other good AI conversations like this?
Yes, the AI models are so useful in areas where you are an expert and can guide it in a useful way. I use multiple chats delay to help figure out client projects.
I think the questions are sufficiently suggestive to induce hallucinations. Still a lot of manual work to check all these claims.
Well said! After a few years of LLM sycophancy convincing laypeople they've managed to make progress on or fully crack famous math problems it's satisfying to see a world-class mathematician really put it through its paces. It's all the more impressive that Tao seems satisfied with the conversation.
This is the second ChatGPT shared conversation I've seen today that is truly fascinating.

The first one was someone proving another conjecture false by just repeatedly saying "keep going" to ChatGPT: https://x.com/DmitryRybin1/status/2079904005652893709

What a world we live in.

Terrance Tao's chatgpt conversation is really interesting for a variety of reasons:

1. The counter example wasn't just a brute force selection, the polynomial is structured in a very specific way that ends up getting the result.

2. Terry Tao's questions are very specific and prompts the AI in a useful way, that without high math training you are not going to get the same information out of it. Terry seems to see some aspects of the problem and counter example and uses AI to brute force some parts of it.

It's crazy how he suggests simplifications over and over and gets led through the finding. Absolutely bonkers how you can use AI to understand something and map it to your own mental map so efficiently, and of course he's most interested in generalizing or finding a simpler sub-result that would explain it.

Just awesome to see new knowledge hit an incredible mind like this. Having these "what if" discussions is what I miss most from JPL and academia.

Math has some of the most insanely dense and impenetrable nomenclature. I can generally keep my head mostly above water or at least near the surface reading from most STEM fields, perhaps leaning on google/wikipedia a bit, but man, mathematics just so quickly decouples from all common tractable understanding it's insane.

Sorry it's a bit of an aside, but I imagine many other otherwise "technical" folks feel the same unfamiliar sense of total loss like when encountering hard mathematics.

What was most remarkable to me from this transcript, was how strong of an equal the AI agent comes across compared to the user (Tao). And Tao is one of the top mathematicians of modern times.

Yes, Tao is guiding it to where he wants to go. But also, Tao is actively learning from it and relying on its explaining, analysis, and inference abilities. You can easily imagine this conversation having taken place between Tao and a PhD thesis student, or even another professor, explaining their results.

What can we imagine and predict about the future anymore? Maybe a year - or two model releases - from now, the AI assistant will be undeniably stronger than Tao, and not an equal anymore.

Jeez. While I obviously can't talk at all about the math, I've noticed a few things:

a) The model thinks on some questions while straight answers on others. (I wish I'd knew from the questions if this is somehow correlated to hard tasks or "inventive" tasks, but that's way out of my league).

b) The model sometimes pushes back. Again, I'd wish I knew if it was warranted, but I counted 2 instances where it said "yes, but with caveats", one where it said "mostly yes but with this correction" and one where it said "careful here, because x y z".

c) The model did q&a + pdf ingestion + code writing + more q&a + thinking + more q&a, for a looong while, while seemingly staying on topic (at least Terrence Tao seems to think they're still productive, so I'll trust that).

This is what model progress is, not number goes up on xBency or yBencher. Damn.

This was my conversation with ChatGPT 4 years ago: https://i.imgur.com/WPaWgzZ.png

Where will we be in another 4 years? What a time to be alive!

"I’ve activated Pro. Can you continue to look for a potential geometric explanation of the X_3 ~ A3 miracle that avoids coordinates or other unmotivated constructions ?"

Another satisfied customer!

It's reassuring to know that even a supergenius's ChatGPT session is one sentence from the human followed by 3 pages of LLM output.
This is the original blog post where Terrence explained his thoughts where the ChatGPT conversation was originally referred from:

https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the...

Similar to how Cypher puts it: I know this is “just” next token inference, matrix mult and just software, ie there’s no “intelligence” there BUT, looking at this convo … damn!

The fascinating this is that the LLM is not acting as a tool here AFAIk, but very much like a colleague.

I have no knowledge of the domain and have only PhD EE level math knowledge, so maybe my bar is too low.

I'll have a blog post up tomorrow about it but the Jacobian Conjecture counterexample is a very funny cognitohazard for LLM assistants. It's a paradox for modern LLMs: they have enough math skills such that they can easily compute the Jacobian to formally verify the counterargument, but its own knowledge base is locked prior July 19th 2026 where all it knows is that the Jacobian Conjecture is unsolved and a random chat user providing such a proof is highly unlikely.
Similar to the story of George Dantzig, who was late to class and solved two open problems in statistics because he mistook them for homework, I think the current batch of frontier LLMs are chained up by knowing which problems are supposed to be unsolved. If they're let free (probably via some targeted RLHF) we might get a flurry of solutions to open problems.
This has received many upvotes and most focus on the interaction between human and machine, which is by all means interesting.

I wonder how many ppl can actually follow what's happening, I mean the math.

How long has it been since we last saw a "LLMs can't really think/be useful/be better than a human expert" discussion on HN? There used to be so many!
I'm watching how Tao uses AI, and it's interesting.

Expand the entire expression, then change the representation to find the core axis. You can't see the axis from just one perspective, so you change the representation. In programming terms, it's like applying multiple domain models. Then break it down into small contract units. Why is it a Jacobian monomial? Why does x satisfy a cubic equation? And so on.

Then swap out the modeling under a hypothesis, assemble it all back together, and verify it through the equation.

This feels similar to modeling in programming.

Observe the whole -> explore better modeling -> decompose local problem -> verify independently -> reason about the highre level structure -> integrate back into the original problem.

This feels similar to when I receive work from a client and write a programming proposal

I don't understand any of the math here, but I had two thoughts. Soon we'll have explainer agents that translate these according to my level so I can, with effort and interest, follow along and stretch my understanding boundary bit by bit.

Two, at some point AIs will be able to use other context like the fact that this is Terrence Tao and not your average Joe and change how it answers, either in tone or structure.

I’ve had a similar experience using LLMs to have mini personal breakthroughs.

One thing I notice is many models say statements along the lines of “okay we have exhausted this thread it’s diminishing returns from here and we should stop and move on”

It’s funny because I’ve been building a tiny neural network maze solver (23 bytes solves 92.75% of unseen 2D mazes)

When I asked ChatGPT/Fable if we had anymore threads to pull to increase capability and decrease byte size, they both basically said no way - back when I was at ~166 byte models with a ~85% solve rate.

Throughout the experiment I just kept trying different approaches and eventually had 3 mini “breakthroughs” in this particular niche. But if I had listened to the models…

Anyway, these models are amazing to experiment with quickly, but they are dumb as hell and so absolute

It's astonishing to me that a technical conversation at that level can be had with a machine that is performing computations to statistically determine what tokens it should use based on numerical weights and relationships. It just doesn't seem like it should be able to come up with details and insights like this, assisting one of the most intelligent humans on the planet to work through something so detailed.
Specifically, if the eigenvalues all have real parts that are negative, then the system is stable near the stationary point. If any eigenvalue has a real part that is positive, then the point is unstable. If the largest real part of the eigenvalues is zero, then the Jacobian matrix does not allow for an evaluation of the stability.

Yes—for a continuous-time autonomous system

x ˙ =f(x),f(x ∗ )=0,

this is the standard linearization criterion, with J=Df(x ∗ ):

If every eigenvalue of J has strictly negative real part, then x ∗ is locally exponentially asymptotically stable. If at least one eigenvalue has strictly positive real part, then x ∗ is unstable. If no eigenvalue has positive real part but at least one has real part 0, linearization is generally inconclusive. Nonlinear terms or a center-manifold analysis are needed.

The last case really can go either way. For example, all three scalar equations below have Jacobian J=0 at x=0:

x ˙ =−x 3 , x ˙ =x 3 , x ˙ =0.

Yet 0 is respectively asymptotically stable, unstable, and neutrally stable.

A slightly more precise wording is therefore:

If the spectral abscissa

α(J)= λ∈σ(J) max

Reλ

is negative, the equilibrium is locally exponentially stable. If α(J)>0, it is unstable. If α(J)=0, the Jacobian test is inconclusive.

This criterion concerns the Jacobian matrix of a dynamical system at an equilibrium; it is unrelated to the “constant Jacobian determinant” condition in the Jacobian conjecture.

Are any of ChatGPT's statements erroneous? (how come Terence Tao gets this high quality output but I get hallucinations when I ask simpler math stuff?)
Presumably this was Sol on xhigh, then over to Pro (as per his indication on chat)?

Is there any way to tell a conversation's model and thinking level?

ChatGPT: "The determinant identity is almost embarrassingly simple once one writes the map in the right way."
I find it amazing how people can use AI to do things that seem hard but yesterday I could not figure out how to install a package on my system. It kept suggesting dependencies that don't exist, and telling me to use functions that are not in the system. The math does not math...
LLMs do match the level of sophistication of the user and the quality of questions they ask. So questions about, how do you improve the level of question? You see the answer in Terry's exchange, both AI and humans are doing an exchange that is thought provoking on both sides for what interesting questions to ask. That is the most startling thing about this -- they are exploring the solution space with the same naivete, to some degree.

Physics, they describe general relativity as a two way street. Space tells mass how to move, and mass tells space how to bend.

Here, Terence tells AI what question to explore, AI tells him what questions to ask next.

The exchange that ensues is just magic to watch.

Would Terrance Tao still be a renowned mathematician if he grew up with ChatGPT?
Thousands of words and piles of equations and not a single mistake pointed out? I have zero domain knowledge here, but enough interactions with AI agents to think this is wildly crazy.
Words and sentences to an LLM are like witchcraft. There are certain words, sentences that make LLMs go a certain way and do vastly better. Sometimes its not at all apparent what set of words will work to do what you want it to do. An example I have been using to do design at a high level is to say to claude.

```

A question is salient to the degree that its answer changes what we do next. Operationally, saliency = the product of four things:

- Decision-leverage — would resolving it one way vs another force a different design or invalidate a stated decision? (No leverage → drop, however interesting.)

- Residual uncertainty given current evidence — is it still genuinely open after reading the docs and the code? (Already settled → drop, however deep.)

- Load-bearing-ness — how much rests on the premise.

- Cost of finding out late — architecture-deciding / expensive-to-unwind raises priority; cheap-to-fix-later lowers it.

```

There are a few things to note about this prompt

1. There is no reason from looking at it that it should work, it even has the word load-bearing which people loathe, but it remarkably produces a stable design with questions from claude (atleast from claude Opus 4.8 and even better from Fable5). Otherwise the design document claude likes to really write are implementation level(code or otherwise). I usually pair this with matt pocock's grilling skill to make claude behave.

2. From design -> implementation, its is generally about understanding when claude is trying to trick you into making something sound like a good/easy solution but has tons of untested assumptions. Here you have to read and patiently spot if a how you would get to the solution is not clear. A common error here are when claude makes a big deal based on what it read and interpreted too seriously without questioning the assumptions. There are several more.

But it also comes down to your experience as a SWE, much like a mathematician's. The frustrating thing about it is, it feels tha a skilled mathematician working with AI can make them productive in ways that are more reliable as compared to a SWE (e.g. lean is deterministic and can provide very strong feedback and LLMs are very good at using that feedback). Maybe a mathematician can chime in on that?

I recall several mathematicians (possibly including Terence Tao) mentioning that fields in mathematics have become so specialized and isolated that a conference like the ICM feels more like a collection of mini-conferences. An expert in one area can barely understand a talk in another.

Modern AI feels like a godsend to mathematicians. It helps them break down boundaries and connect concepts in ways a mere mortal couldn't imagine.

The biggest difference I see between the way I use LLMs and Terrence Tao's is that he's not constantly swearing at it.
The most important prompt lesson everyone should take away: never ask yes / no / leading questions.

Almost all of Tao's questions begin with what, or why. That forces an open ended response, which is a great way to reduce or eliminate sycophancy and severe hallucinations. The less you steer, the more accurate it gets.

It's fun if you ask ChatGPT to guess the identity of its interlocutor :) It will guess math researcher or paper author without hints, but if you give it some additional hints, "this was shared over the internet", "it's someone willing to work with AI", it will guess Terence Tao as the first choice.
Seems important that the breakthrough conversation with Fable that led to the counter example should be shared too.
Hey, Mr. Tao, where would you put the intelligence of the current models compared to top humans you surely mush have worked with? Perhaps given in the count of people you know that are higher than the newest LLMs? Also, how would you rate the speed of work compared to what your speed is, if this makes any sense?
Imagine the AI's of the next century trying to explain incomprehensible dark matter math to the smartest humans of the time. David Deutsch says the human mind can comprehend _anything_. I think the next decades of AI-driven math research might put that assumption to the test.
Not sure how to suggest edits to titles, but @dang: his first name is Terence, not Terrence.
We had this yesterday (317 points, 133 comments)https://news.ycombinator.com/item?id=48998362

Classic Gabe-bot

Can't even ctrl+f the conversation, wish openai would fix that