back

by rvz·13d ago·view on hn ↗
What does this even test for? Can I use LLMs to directly generate machine code to replace my compiler? Or maybe I can use LLMs as a bare metal OS / scheduler to replace my machine's operating system and scheduler? It makes zero sense to test for that.

Not only that this so-called "benchmark" isn't economically useful, but that it tests for the sake of testing; and for attention.

To end this obsession with generating SVG pelicans, Quiver AI's [0] model is actually designed to generate SVGs from prompts and has done so for years.

So there is no need to continue with this un-serious pseudoscientific "benchmark".

[0] https://quiver.ai/

1 comments
It tests for the "I" in "AI".
So if we test models that only output text to directly generate waveforms of sound from text or binary code, or directly generating binary code to replace a compiler, does that mean it is "intelligent"?

Does that even test for intelligence?

This is like testing if a horse can fly just because someone showed an image of a Pegasus, or testing if a fish can climb up a tree and believing they are not intelligent because each of them cannot fly or climb up trees.

I'm not sure where you are coming from, but we are testing a model that can create images to create an image for us. If it can't even do that well then I'm not sure why we need to talk about horses here.

Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after. (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)

> I'm not sure where you are coming from, but we are testing a model that can create images to create an image for us.

Then use a model that supports multi-modal output then, instead of using models that only support raw text output to construct an image.

> If it can't even do that well then I'm not sure why we need to talk about horses here.

Testing a model that only outputs text and forcing it to output an unsupported modality is obviously not going to work well which is what it going on here.

Would you expect Claude/GPT to replace your compiler because it can output text? That would be like expecting that horses can fly because someone saw an ancient winged-horse on a brick tablet.

> Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after.

That is because the premise of the test is ridiculous and pseudoscientific.

> (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)

Why did you say it was a "test for intelligence", when we both know it is a joke that tests for nothing?

One way to characterize intelligence is the capability of solving problems one has not specifically been trained for.

Somebody who has trained painting animals and vehicles for decades is surely skilled at that, but is not solving new problems, so is not necessarily applying intelligence when painting something of that sort.

You can take a reasonably skilled programmer, show them some C code and ask them to write equivalent ASM code, perhaps providing them with a language reference for either, and they will be able to do this translation. Slowly, but they will get there. They have not been trained to do this, but they will "figure it out". With enough time and motivation, even non-programmers would be able to do that. That is what intelligence is and of course Claude/GPT needs to be able to do that if it wants to claim that it intelligent. No flying horses.