I think the assumption that there is some sort of formal ‘reasoning’ process that is inherently superior to whatever ‘non-reason-based’ decision process ‘non-reasoning’ systems use is not immediately justifiable.
Most reasoning is ultimately based in heuristics: this is the right thing to do next because it has worked in the past. This is the right thing to do because our intuition tells us doing it will lead to this desirable outcome.
When you’re formulating a math proof, sure, the steps you are laying out make up a ‘reasoned’ argument but how you choose what steps to take is wholly intuitive. Why assume the contrapositive to begin this proof? It’s worked in the past and it might work here.
It’s all pattern matching trained on a reward function.
At the end of the day, there's a difference between showing your work for a math problem by writing down each logical step as you work through it, vs. just blindly copying both the answer and the "show your work" part from your classmates.
That humans can fail at math, or at other kinds of logic problems, is immaterial. It's still an entirely different process, and it leads to different results.
The fact that humans often fail because they're actively prioritizing something else can also be for the best. Like, it's reasonable to get a math problem wrong because you're distracted by argueing classmates who might be about to start a fight right next to you. Or you might realize that the premise of a question itself is wrong. Or you might logically know that the Earth goes around the sun, but decide to pretend otherwise because it's not worth dying over.
Humans are always balancing priorities, and it doesn't always make sense to judge their success by a single metric like "do these math problems and show your work within 10 minutes". LLMs do not have these conflicts, even when we would want them to.
I suspect humans have other ways that help with error correction and guiding the reasoning effort, but that's another story.
Surely it is worthwhile to attempt to understand the details of that? And if we seek human equivalent performance then it is reasonable to wonder if the reasoning achieved to date is the "correct" sort.
Why pursue this one avenue vs some other, typically its cause it seemed more promising, and the person can come up with reasons but did they explicitly verbalize a fully sound chain of thinking at the time? Probably not.
Not to say that the explicit thinking, or writing things down, isn't important. But if I examine the process by which I develop a proof or something, there's a lot of vague hunches, blind alleys, etc that come along the way. And, many of the blind alleys etc probably aren't actually that important in the end for me finding the right answer -- if you were to trim that part out of my own internal reasoning trace but left the rest intact, I'd still get the right answer because, well, it was a blind alley.