back

by padolsey·9mo ago·view on hn ↗
I help run an eval platform and thought it fun to try a bunch of models on this challenge [1].

There's some fun little ones in there. I've not idea what Llama 405B is doing. Qwen 30B A3B is the only one that cutely starts on the landscaping and background. Mistral Large & Nemo are just convinced that front shot is better than portrait. Also interesting to observe varying temperatures.

I feel like this SVG challenge is a pretty good threshold to meet before we start to get too impressed by ARC AGI wins.

[1] https://weval.org/analysis/visual__pelican/f141a8500de7f37f/...

2 comments
> I feel like this SVG challenge is a pretty good threshold to meet before we start to get too impressed by ARC AGI wins.

It's a very bad threshold. The models write the plain SVG without looking at the final image. Humans would be awful at it and you would mistakenly conclude that they aren't general intelligences.

I dunno. A competent human can hold the a mental image and work through it. Not too hard with experience. What I generally mean tho is: I don't think we can state the supreme capabilities of AI (which people love to do with grate fervour and rhetoric) until they can at the very least draw basic objects in well-known declarative languages. And while it may be unwise to judge an AI based on its ability to count the number of 'R' letters in various words, it -- amongst a wider suite ofc -- remains a good minimum threshold of capability.
You will never guess what the upcoming gemini 3 has been benchmaxxed upon ;)