back
139 comments
When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.

Are we missing the forest for the trees here? If a math problem falls in the forest but nobody is around to understand it does it make a sound?

How can we possibly make use of these breakthroughs if we don’t understand them? How could we ever make anything useful with them?

Are we ready to just let go of our intellectual faculties and give them to a giant supercomputer nobody understands? How do we tell truth from fiction?

I get the impression that the value of unproven conjectures is more in the new math and techniques that may be discovered - by humans - trying to prove/disprove them, rather than much utility in any eventual result.

Take something like Fermat's last theorem - I'd be curious to hear of any use of the result itself, but there was a massive amount of new mathematics generated by those working on it, whether ultimately successful or not.

These AI math proofs are interesting testament to the power of reinforcement learning applied to math, obviously reflecting the axiomatic self-consistent nature of math itself, but it doesn't seem they have the same value as a humans working on these problems since they are using known math to solve them rather than inventing anything new.

However, it would still be interesting to analyze the LLM lines of reasoning that lead to any of these results, since there may be value there even if no new math, just as human Go players have found value in analyzing computer Go.

Still, as Demis Hassabis has himself said, the real goal with AI is discovery and creativity - you want to create the thing that could design the game of Go in the first place, not just play it. Similarly with math, while there is interest in seeing an AI "play math" using the rules of the game, what would be of much more interest is the AI that can create new math, in the same way as Andrew Wiles did while proving Fermat's last theorem.

> When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.

Math is an incredibly broad field. I mean, you don't expect a traffic engineer to understand anything about nuclear reactors, do you? Yet, they are all 'career engineers'.

This seems like an extreme exaggeration from a few people claiming to not understand a very recent result. This cycle of a new result being discovered, and reaearchers needing some time to truly digest and disassemble it, is normal.
I think it's the scientist version of "I vibecoded ten apps this weekend (at one point I'll have real users too)".

Because it's math, it's all mysterious and genuinely impressive, but in the end, if no human cares about it (apart from attention grabbing "it's so over" tweets and articles), does it really matter?

Latest presentation of Terence Tao on what current advancements in AI mean for math discusses (among other things) those issues: https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.p...
The field has been dealing with this for a long time.

Since at least 2014, which was my first brush with the phenomenon when someone published a 13GB proof [0].

The consensus is that such a proof is potentially illuminating, though further work is likely required. If for instance, conjecture A is true if and only if conjectures B & C are true, and B is proven false through one of these such proofs, then we can see that A is also false given that we accept the disproof of B.

Though, the sense is that further work is likely required because it is easy to see that further work along the same direction, or in directions depending on the proof will be hard or impossible if there are not enough humans or agents that are capable of understanding and utilizing the proof. Making it 'more elegant' will increase it's utility despite not proving anything new.

This is adjacent to all of the work done to create multiple proofs using different techniques. Having the same information (that X is so) in different languages (algebraic, geometric, via harmonic analysis, etc) allows for researchers not familiar with the original technique to participate in further research.

[0] https://www.newscientist.com/article/1997488-wikipedia-size-...

Mathematics is an unusually dense (if not the most dense...by a few large steps) field. So lots of areas of mathematics are extremely deep and narrow without any real shortcuts, even for seasoned mathematicians.
I don't mean to be dismissive, are these just old puzzles with no practical use whatsoever?
The field of mathematics is smart about this and knows the difference between a pile of Lean code and understanding, and mathematicians try to get from the unintuitive explanations to something that makes more sense, e.g. https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the... (where incidentally Tao used a chatbot to help take apart the problem, but with a lot of interaction and work from his side).

Making things understandable is mathematics, and more generally a kind of intelligence, and is crucial to continued progress. You couldn't use algebraic geometry to disprove a conjecture if people hadn't organized (what could have been just) a pile of random observations into something called algebraic geometry.

Historically LLMs have done best where it's possible to train using an objectively verifiable reward function. Computer programs are pretty good on this front and so are Lean proofs. (Of course, they don't only do things you can RLVR heavily, but those have progressed fastest.) Not sure where 'making mathematical knowledge more understandable' falls on that spectrum.

Understandability isn't only important for advanced math. Keeping computer programs from becoming a mess is a challenge in high-level organization too, and the chat with the user is an explanation task. If you look online at what people say about large LLM-built codebases (SlopCodeBench is a neat effort to make make it concrete, but common wisdom seems to mostly agree on the general problem) and chatbot prose, I don't think everyone considers those solved problems!

It's hard to tell how thoroughly the labs grasp and care about this at an organization-wide level. I'm sure at least some maybe-results exist inside labs but haven't been published because the humans couldn't verify them and didn't want to be embarrassed with a false result. (Maybe also why counterexamples are a lot of the first results published: often simple to verify, even if hard to obtain.) A good sign would be if results in a few months come out more like what mathematicians consider well-written papers explaining results in a more intuitive way, fewer shocking announcements of bare counterexamples in tweets. It's probably a slow climb to get there.

This has been an ongoing debate since the computerized proof of the four color theorem fifty years ago.
Your observation is perfectly on point, I think the season of companies announcing breakthroughs might be over soon (unless they somehow manage a major achievement, P vs NP or similar). At the same time, mathematicians will be left with superintelligent machines solving the actual math for them, much like software engineers nowadays. This was unexpected, and unexpected at this scale up to a couple of months ago.
> Are we ready to just let go of our intellectual faculties and give them to a giant supercomputer nobody understands

An unsettling number of people do seem to be more than ready for this. In fact I'd say they seem almost gleeful about it. Thinking is hard and they don't seem to like doing it!

> How can we possibly make use of these breakthroughs if we don’t understand them? How could we ever make anything useful with them?

Theoretical science isn't about usefulness per se, but knowledge and understanding for its own sake, so a better way to say this is to point out that you only benefit in such cases if you understand things yourself. If understanding is the goal - which it is in the case of true theory - then the only way to attain that goal is to actually attain understanding. What good is it if an LLM produces a valid proof, but no one grasps it? This isn't like digging a ditch where it doesn't matter who does it or how he does it as long as you have a ditch. Here, the ditch is knowledge, as it were.

My perception is different. I start from the axiom that AI is a massive compressed corpus of knowledge. That it can find solutions suggests to me that the solutions were already known, but simply lacked publicity. This is less about discovery and more about pattern-matching.
I find some interesting parallels in chess, which often has lots of analogs with math to begin with. But chess went from a pure human endeavor to one where supercomputers aided by world class players finally managed to eek out a slightly suspicious win against a world champion (approximately where we are now in math) and to now a days - where your phone could easily crush the world's strongest player, who is also probably the strongest player of all time.

The way the chess world adapted this was initially to try to understand the machine. After all chess, like math, is complete information - so you can easily see the computers 'thoughts' in terms of the exact moves its saying are best in a variation and how it might respond to any other idea. But it quickly became clear that this wasn't working so well.

Players would regularly get positions that the computer says 'and black wins' and then proceed to lose it convincingly, simply because the positions were so extremely weird and difficult to play that even if it might be technically winning, it's the sort of position where you're walking a fine line with lots of complex moves to find. Humans aren't computers and even the best of us can't play like one in weird positions.

Now a days they're taken more in balance. The computer's evaluation of a position is probably about as good as you can get, but playability matters much more in practical terms. Knowing the eval of a position doesn't really matter if you don't understand the position. Knowing the answer can help with understanding (for instance computers have radically reshaped and improved human understanding of space in chess as we noticed computers obsessing over it) but I think the days of 'oh the computer says it's winning, so I should be able to take it from here' are near to gone.

It's more like FOSS project using and you ask a developer who works at a different company how to build this project and he says "I don't know automake, perl and there are lots of macros"

What are the odds that this random FOSS project has solved the problem of building software and nobody noticed. Close to 0. Don't confuse accidental complexity for transcendental depth.

Even the experts on the problem being solved find the writeups nearly impossible to read.

Example: https://nitter.poast.org/henryquantum/status/208362369543662...

Seems like a disservice to the community that openai put so little effort into producing good writeups...

> How can we possibly make use of these breakthroughs if we don’t understand them? How could we ever make anything useful with them?

Isn't that what people from more practical sciences said about math anyways?

All math is eventually applied math.

In college the punchline for all the engineer, physicist and mathematician jokes were something like "The mathematician says: Yes, there is a solution."

In all serious "I don't understand any of this it's way over my head."

Maybe the "beautiful, elegant" math is really just accidentally that way, just the tiny cross section our dumb human brains can understand. The vast majority of it could be inscrutable, ugly, chaotic and seemingly meaningless.
> When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.

Huh?

Mathematics is an extremely wide subject, it's perfectly normal even for two professional mathematicians not to understand each other's work. Have you considered that maybe they just don't work in that area?

Are you implying OpenAI's paper (which was, by the way, edited by humans and provided Lean certificates for most of the proofs) is actually gibberish? That's flat-earth levels of conspiracy.

i doubt anyone really wants, or cares, to understand what the AI comes up with. If it's not related to your work or your name isn't associated with the discovery then I would think time is better spent on things that are.
They claimed them as advances, not breakthroughs.
There's no honor in being stupid, yet I must admit I am stupid, as there's even less honor in being stupid and pretending otherwise.

A lot of people were super hyped about OpenAI's 10 discoveries, but I still don't understand what they mean, and even if I did, what are the implications.

Like, what are non-sofic groups, and what follows from the conclusion that they exist?

I mean in the sense that quantum mechanics might make my head spin, but it's because of that that we have stuff like semiconductors, which have been one of the most significant discoveries.

The Fourier transform is one of the reasons we have fast telecommunications and radars.

What practical things are possible or might be possible due to these results?

I have the same questions about human mathematicians!

I can't tell if xkcd #435 is still true, or if math is just as mushy as everything else seems to be. When a math proof can only be understood by a handful of people, what does that mean about that proof? I think the LLMs are pushing a problem that existed already and pushing it further.

[0] https://xkcd.com/435/

>If a math problem falls in the forest but nobody is around to understand it does it make a sound?

This is not new nor unique. There is plenty of research, especially in math, which can really only be understood by a few people in the entire world. It is not uncommon for a proof to be presented by a mathematician which, initially, is only understood to that mathematician, and it can take a long time for even another mathematician who is an expert in the same field to be able to confidentially say they understood it.

>I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.

This is meaningless in a vacuum. If you give a novel proof in some niche subfield of topology to a competent mathematics researcher who focuses in number theory, they'd say the same thing regardless of if a human or machine wrote the proof. A bunch of "career mathematicians" on Twitter proclaiming this doesn't mean anything other than these people aren't currently equipped to understand the contents of the proofs. That's fine and normal, but the idea that anybody with a PhD in Math should be able to pick up one of these proofs and give it a skim and be able to say, "ahh, yes, quite clever, it seems so obvious in retrospect," is absurd. That's not how this kind of research works.

>Are we ready to just let go of our intellectual faculties and give them to a giant supercomputer nobody understands?

Nobody is blindly accepting these proofs as valid. ChatGPT isn't spitting out a wall of text and proclaiming that they've solved a previously unsolved math problem while everyone is saying, "well if an LLM says it, it must be true!" lol

These proofs are being checked by automated systems (which have been in-use well before LLMs have existed) as well as being checked over by actual experts who are actually capable of (and motivated to) verifying these proofs. But that work still isn't done. There's enough evidence that these companies are confident in saying these proofs are correct, but there's going to be a lot of ongoing work from people to continue to verify and, more importantly, understand these proofs. It's literally some of these people's full-time jobs to do this.

>How do we tell truth from fiction?

When was the last time you verified even a classical, relatively simple mathematical assertion? How often are you just relying on a larger system of experts to ensure that we're not just blindly accepting fiction as truth?

That's not to try to stick it to you personally, but it's just highlight that there's an entire system in-place here that you're not aware of and that you don't have an understanding of that is working just fine including in this context. Real mathematicians aren't going to lazily start letting OpenAI assert whatever they want about their products solving these kinds of problems without heavy scrutiny.

> How can we possibly make use of these breakthroughs if we don’t understand them?

I mean, arguably the development of LLMs matches that description.

I, for one, welcome our giant supercomputer overlords.
Non-Erdős problems are also falling.

The AIs seem to have some combination of very broad familiarity with math (enabling relevant things from other subfields to be brought in to the proof) as well as patience and "sitzfleisch" (stamina in working through details even if they aren't immediately obviously promising.)

An obvious area for improvement would be automated generation of new conjectures and attempts to prove (or disprove) them, with the discovered arguments then being used as training for refined models. This will require autoformalization to check the results as there will be too many for manual verification.

Generating new conjectures is easy. It's generating novel conjectures that's hard.
The mathematics are above my intellectual capacities, but I still find the man and his lifestyle fascinating. Though I admit his heart attack probably wasn't a coincidence, sadly.

I wonder if national, institutional, or otherwise "eccentric" sponsorships (encouraging a similar migrant-madman approach to academic cultivation) of some of the folks on HN wouldn't lead to meaningful discoveries in CS.

I often see comments that some of the users here long to "make a computer do neat tricks all day", and I can't help but think meager sponsorship could go a long way in this area. Existing grant structures, being much more traditional, are constrained by their cost.

> I wonder if national, institutional, or otherwise "eccentric" sponsorships (encouraging a similar migrant-madman approach to academic cultivation) of some of the folks on HN wouldn't lead to meaningful discoveries in CS.

The problem is that Paul Erdos was eccentric but he was also Paul Erdos. I'm not saying you're implying that, but I feel it's a similar line of thinking to how popular culture often romanticizes autism and Asperger's because some very smart people are (allegedly) affected. The same group includes people who need 24/7 care.

Computer science at the frontier is as specialized and hard as mathematics, you need years of study to truly understand a field well enough to make meaningful contributions. You can try sponsoring me if you really want, but I don't think you'd be spending your money wisely in expectation.

> Though I admit his heart attack probably wasn't a coincidence, sadly.

He died aged 83...

When Gwern needed a fellowship to come to SF he got one. When he had his “great idea” it was funded. The truly unique people out there find a way. Any government bureaucrat will just find the Jason Ardays of the world.
Honestly, I think our extremely liberal disability benefits scheme achieves this to a large extent (albeit possibly with downsides for both the recipient and society).

Look at all of the crazy stuff made by unemployed trans people/autists over the past decade, disproportionately from countries with strong welfare states.

something i've just realized : today long-standing maths problems are falling. It's great intellectually but won't probably have an immediate impact on our lives.

Now, what will happen once long-standing physics ( and chemistry and biology) problems will start to fall and at the same rate ?

Then we're going to enter a totally different world.

It's hard to see problems in those fields falling at anywhere the same rate as math, because they are all experimental fields.

There may be some problems of type type "why does X happen?" that appear answerable in terms of known science, but even these would need verification. If you want to make advances in fundamental physics, then a promising AI-generated theory might take a decade and billions of dollars to prove or disprove.

Math is a rather unique field in being entirely theoretical, axiomatic and self-referential. It is basically the best possible case not just for AI to advance without needing experimental verification, but also specifically for today's AI technology of auto-regressive LLMs and RL training, whereby valid reasoning steps learnt in one context will also be valid in another context (i.e. there is some generalizability of learnt reasoning) as long as you have learnt the pertinent aspects of that context that the validity depends on.

For example, protein folding has been figured out by AI. This was a very big moment for science and yielded a nobel prize. Alphafold 1 happened 4 years before the first version of ChatGPT and the LLM craze we see today.

But yeah, most problems in physics, chemistry or biology require labs on top of actual hard thinking. You need to be able to design experiments in a certain way. Once you have the funding, the right tools, the right people to use those tools, then you can use LLMs to increase the speed of the calculations and so on.

There have been other discoveries though by deepmind: https://deepmind.google/blog/millions-of-new-materials-disco...

Those long standing math problems solution may serve as building block for solving the experimental science ones.
Disappointing lack of "Why" in an article that starts with it.

Are they actually doing something new and novel, or are they just absorbing that "a=b as was proven in transcendental hyper-circular group theory; and b=c was proven in universal quantum superposition"; and they're the first to find the connection that a=c? And several of the problems are counterexamples, not novel proofs of correctness?

Its fascinating either way, but it'd be nice to actually understand more of what is happening.

It seems that AIs are really good at finding counterexamples now.

Even if progress by AIs in proving conjectures lags, it seems likely that AIs collectively will, in the next few years, find counterexamples to nearly all the Erdős (and other) conjectures that are actually false and also provably false.

That means we will able to assume that nearly all the remaining conjectures are either true or undecidable.

Surely, that's good for folks who just want to know where the truth boundaries in mathematics lie.

It's obviously causing a lot of soul-searching amongst professional mathematicians.

Arguably, they should have given less weight for the last 100 years to Hardy's view in 'A Mathematician's Apology' [0]:

> It is a melancholy experience for a professional mathematician to find himself writing about mathematics. The function of a mathematician is to do something, to prove new theorems, to add to mathematics, and not to talk about what he or other mathematicians have done.

Rota takes a much more balanced view in 'Indiscrete Thoughts' [1].

"Problem Solvers" take Hardy's view:

> ... The mathematical concepts required to state mathematical problems are tacitly assumed to be eternal and immutable. Mathematical exposition is regarded as an inferior undertaking. ...

While for "theorizers":

> Mathematical exposition is considered a more difficult undertaking than mathematical research.

If professional mathematicians can reinvent themselves, there will be plenty of work left to do to explain the results of AIs to other humans.

There probably needs to be a new career path into professional pure mathematics other than doing novel research in a PhD.

[0] https://en.wikipedia.org/wiki/A_Mathematician%27s_Apology

[1] https://ncatlab.org/nlab/show/Gian-Carlo+Rota

https://ncatlab.org/nlab/show/Gian-Carlo+Rota

Because pattern recognition and deduction is automatable and LLMs are superhuman at low depth high depth problems. Next
is probably a mix of other math papers that combined can solve this, the ingredients were already out there and they got trained with it.
> “A big problem is AI is being used a lot by people who aren’t mathematicians, who don’t have a huge mathematical background and are not capable of verifying the output,” Bloom said. “They like to move fast, ask their AI to check it, it grows and grows. We’re seeing a lot more of these 100- to 200-page papers that people are posting. ‘I solved this theorem; I got AI to generate the proof and check the proof and write the paper.’ But no human has read it, and no human is going to read it. It’s a huge challenge now.”

Mathematics has the same problem as open source projects that are overwhelmed with AI slop!

Quantamagazine is owned by the heavily AI invested Renaissance Fund.

The whole article is an ad that covertly or overtly inserts how websites are built with ChatGPT, how humans say that AI is better than them etc.

This is incidentally the future of chatbots. I could not have written this comment without Illy Espresso. Would you like to find a cafe near you?