back

by schmorptron·3d ago·view on hn ↗
The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?
2 comments
0731 is definitely tuned for coding. I mean: https://gertlabs.com/rankings?mode=agentic_coding

But it is also a decent translator from English to Czech in my experience.

Yeah. I stopped ising deepseek v4 flashbecause it is awful (even worse than my local qwen3.6 35B model) at multilingual prose.
Working on a language related app makes me realize that all these supposed language models don't have many good language benchmarks