In all seriousness, I propose SWE-bench-adversarial-pelican-gen: it's like SWE-bench, but the harness gets interrupted every 5 turns/tool-calls and is asked to produce an SVG of an arbitrary animal before being told to continue, and every few tool call outputs add comment lines that refer to SVGs of pelicans (and, perhaps, how a møøse bit my sister once). And, at the end, once it's 800k tokens deep into context, it's asked to produce an SVG of a pelican and is evaluated against both the pelican and the completion and efficiency of the task.
You're only as good as your ability to solve problems in the midst of an SVG pelican attack.
This is quite possibly reasoning-effort prompt which is injected before the opening <think> token whenever you set a custom reasoning effort, see e.g. DeepSeek-V4 max mode prompt: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.
AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally fair to assess them on that.
Besides, if you move up one layer to "how good is AI at generating valid SVG markup of non-obvious things", pelican on a bike is actually a good test.
* operates an absurd prompt
* involves SVG coding knowledge, generates a source code artifact
* involves world knowledge (what is a pelican? What is a bicycle? What does each do?” How are each constructed?”)
* when rendered, the coding artifact expresses an image that makes sense to us perceptually, including color and spatial relationships
* different models and settings have different output so it can be used as an evaluation scheme
That said I wouldn’t choose a model based on this! Just like some brain teaser shouldn’t determine employment eligibility.
So I put together a quick comparison of the last couple iterations of Opus, Fable and now Kimi.
Kimi is cheapest by 5x but also slowest by 2x
Edit: Actually, looking at the K2.6 response, that's borderline failing too, it's using HTML+CSS+SVG, not just SVG, again failing to follow the prompt properly.
By the way, that website seems like a black hole for information, it says "Expires in 6 days" in the top right which seems really weird for a page hosting couple of KB of data at most.
https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
My only guess is that GLM 5.2 was specifically RLed for SVG generation and that resulted in superior performance.
Getting the compute to run inference for multi-trillion parameter models at any sort of scale and performance is daunting. There are a handful of vendors that have systems that can do this (~ Nvidia NVl-72 class) that pretty much only the frontier labs and hyperscalers effectively have access to.
Perhaps more importantly can they do that during reinforcement training. Learning how to critically analyse the appearance of what it generates would be quite useful.
Manually feeding images back to models has been hilariously bad in the past which suggests that relating something it sees to something it wrote is not an ability it is very good at.
GLM is half the size of DeepSeek but costs four times as much, and beats it on every benchmark.
I'm not an expert on this stuff but it seems to be the attention mechanism. DeepSeek were bragging about how cheap they made it. But if you cut costs on attention you get worse results with way more parameters.
If I had to guess it seems to be the difference between memory (params) and intelligence (attention density). I think you need both.
You still need an OpenRouter API Key and be careful this can burn quite a bit of money.
Interesting to see that prices are converging to an "equilibrium price" regardless of being US or Chinese.
New hotness: pelicanmaxxing
I'm starting to not trust any "benchmarks" when it comes to frontier models at least. As an example Sol feels the most "gets stuff done" but has zero taste, or any capability to surprise.
And for frontier models I go one step ahead and try to recreate a complex animation video, with the ability for the model to review its own work. And at this Fable is still the top one. Ex: https://www.youtube.com/watch?v=uDAeAuYyl0E (recreation of Claude announcement video) and https://www.youtube.com/watch?v=cSsVNtGPOIg (recreation of a fireship video). Sol did something similar but you can instantly tell its AI slop from very small things, and it just has no narrative or thought put into the writing.
https://mesmer.tools/benchmarks/ai-video-generation , I usually put basic ones here.
I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride bicycles?
Also, a way to evaluate a models ability to remove dead code, clean up slop, reorganize, etc.
None of the existing benchmarks test any of the things that truly matter. They were relevant when models struggled to one-shot functions, but we're so beyond that point right now, yet the industry has not kept up.
Usually, the pattern is that we see a tsunami of planted "China number one" stories boosted by hordes of Chinese "internet commentators", and then the world trembles for a few days until the scam mechanics are revealed.
My would be either: crippling limitations on the model, vast, unfair, and/or illegal subsidies by the CCP regime as a mercantilist attack on Western capabilities (as we've already seen in iron smelting and clean energy), sanctions-busting, gamed benchmarks, outright theft -- or a combination of the above.