https://files.parasmittal.com/openai_aa_luna_dsflash.svg
1: https://openai.com/index/advancing-the-price-performance-fro...
High hopes for V4 Pro
So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the same model with fireworks/openrouter, with zdr thrown in, token costs ratchet up with no explanation. Likely that the model is subsidized for gathering usage data. I am waiting for the day I can run this locally.
"For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95."
Does the file hosting actually cost peanuts when you do it yourself and the cloud has shattered my understanding of what it actually costs to deliver so much data?
At the scale of Huggingface, that still amounts to a lot load of money. Significantly less than if you did the same in AWS, but still a lot
That said, they do have a deal with AWS to make the data available in AWS ip space. Maybe they got some cheap hosting out of that too
As long as you're smart on caching on the edges and deduplication/overlaying with your content (which structured data types like models, containers and repos typically are fit for) you can get remarkably far for less than you'd think.
The economics of serving models seem to be way more dodgy.
https://artificialanalysis.ai/models/deepseek-v4-flash?intel...
Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.
If the full non-flash model follows up with the expected improvements, and at the price point they've been keeping, it puts the frontier labs in a tough position and it feels to me like like OpenAI is reaching deep into their pockets to try to head that off.
TFA link is a 404 though. I'm reading through https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 instead
It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
It use the stock model, no new models requires.
Worth spend a few hours to try.
The DGX Spark requires a small hack to ignore the difference between sm120 vs sm121, but it does run on sm121.
The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
The specific agent is focused on getting precise and on point answers about a codebase.
The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.
The benchmark included more than 50 questions or different difficulty.
But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.
Just to say that the quality of the harness is as important as agents intelligence.
But a REALLY good pelican on reasoning mode high (via OpenRouter): https://static.simonwillison.net/static/2026/deepseek-flash-...
Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.
Similar price? Doesn't make sense. Maybe they meant power, capability or speed?
[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...