back
5 comments
> So then you need to explain ARC-AGI-3: https://arxiv.org/abs/2603.24621

I don't, we originally had the turing test which was designed to determine human intelligence by its ability to imitate us with natural dialogue, but we've since defeated that. I stated "to me" because it's my personal opinion on a definition whose goalpost will probably never stop being moved.

> Back 1996, EQP automatically solved the Robbins conjecture. But nobody concluded EQP was generally intelligent.

EQP doesn't have the 3 criteria I outlined, which were different than "solving a math problem"

You wrote:

> > It's the cumulative knowledge of all general human intelligence, baked into an artificial form, which can then use that knowledge to achieve novel goals.

But it (currently) can't solve puzzles like ARC-AGI-3 that children can solve.

look up the actual origin and purpose of the turing test, you may be surprised
One question I have about ARC-AGI-3 is how much it depends on vision ability, which has a substantial hardware component that isn't "General" "Intelligence".

What happens if the game is encoded in a non-visual logical form?

ARC-AGI-3 is actually about action efficiency not solve rate, current LLMs just make a bunch of moves that inefficient, thereby lowering their score
OpenAI claims that the harness that ARC used was unfairly handicapping the model: https://openai.com/index/how-two-settings-tripled-our-arc-ag...
I think ARC-AGI-3 specifically forbids harnesses. This means that you're basically limited by the context window, so it's no wonder that LLMs can't do that well. Unofficial versions that use a harness seem to be doing fine on it.