back

by rcarmo·6mo ago·view on hn ↗
I don't think this is a good "benchmark" anymore. It's probably on everyone's training set by now.
2 comments
I think it could still be an interesting benchmark. Like, assuming AI companies are genuinely trying to solve this pelican problem, how well do they solve it? That seems valid, and the assumption here is that the approach they take could generalize, which seems plausible.
The point of this benchmark is that making decent SVG art is actually useful. Simon has private image prompts he uses, since he didnt say gemini failed at those it is reasonable to assume those were also successful.