back
451 comments
Nice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
For a while now, I've found pelican rendering to be an unreliable metric for LLM ability - and most people know it. Yet, somehow it gets upvoted to the very top of every new model discussion.
On either side of the front wheel is a perfectly reasonable place to carry cargo. I think I'd have taken more issue with the spokes, or at least that's what stood out to me. The chain is indeed nice, however.
You don’t need the best model in 99% of cases…
Might be time to also have it try more three dimension pelican rendering. Or a short animated version (even just a few frames) still in standard SVG
these links never work for me. always "Error: Enter a valid URL" when opening in Firefox. maybe a URL escape issue with Glider?
we hitting singularity levels of bicycle chain here
Basket? Fish? All I see is the model recursively running itself locally on an eye-pad, which for some reason beyond our understanding is obscuring the invisible fork.
Your tool is giving "Error: Gist API returned 403"
I think I saw a better overall composition out of Flash 0731

Effort on this one?

Deepseek V4 Flash 0731 was such a massive jump in capability for such a small model (and price), that I'm a bit disappointed by this release.

I keep my agents on tight leashes, using them very interactively for bouncing off ideas, architecture, and then writing code (especially prototyping) and Flash has been crushing everything I ever needed it to do.

Maybe my ambitions are too tame compared to people needing Fable / Sol grade models, but I'm probably staying on Flash and not moving on to Pro for the foreseeable future.

Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one...

Tested this model, and gpt-5.6-terra-high.

Results: this one had few issues. terra: none.

These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different.

They are good till the project is simple... not anymore.

I've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts.

I can't wait to try this new one.

What I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.
Have been letting it spin pretty hard (~$12.50 for 2B, 50% cache hits) on my traffic simulator/distributed physics engine all day, it's found some pretty significant gains without introducing any new problems.

I'm happy

Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project.

Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.

Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.

Based on my experience so far, compared to previous models, DeepSeek V4 Pro achieves results equal to or even better than before, but at a lower cost.
Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
Benchmarks:

    | Benchmark                | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2   | Kimi-K3   | Opus-4.8  | Fable 5       |
    |                          | 0813      | 0731        | Preview   | Preview     |           |           |           | (w/ fallback) |
    |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
    | HLE (wo/w tools)         | 42.7/60.0 | 37.8/51.5   | 37.7/48.2 | 34.8/45.1   | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0     |
    | Terminal Bench 2.1       | 87.9      | 82.7        | 72.1      | 61.8        | 81.0      | 88.3      | 85.0      | 88.0          |
    | NL2Repo                  | 61.5      | 54.2        | 38.5      | 39.4        | 48.9      | -         | 69.7      | -             |
    | Cybergym                 | 83.3      | 76.7        | 52.7      | 38.7        | -         | 80.0      | 78.3      | 83.1          |
    | DeepSWE                  | 62.7      | 54.4        | 12.8      | 7.3         | 46.2      | 67.5      | 58.0      | 70.0          |
    | Toolathlon-Verified      | 74.1      | 70.3        | 55.9      | 49.7        | 59.9      | 76.5      | 76.2      | 77.9          |
    | Agents' Last Exam        | 25.7      | 25.2        | 16.5      | 15.8        | 23.8      | 27.6      | 25.7      | -             |
    | AutomationBench (Public) | 31.8      | 25.1        | 12.8      | 10.8        | 12.9      | 30.8      | 27.2      | 29.1          |
    | DSBench-FullStack        | 71.1      | 68.7        | 41.8      | 37.0        | 61.8      | 73.7      | 71.6      | 77.2          |
    | DSBench-Hard             | 67.2      | 59.6        | 31.1      | 25.8        | 54.5      | 63.0      | 71.7      | 68.3          |
Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...
It appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.
Again, I will wait until there's a provider that doesn't train on prompts before I will benchmark.
https://api-docs.deepseek.com/quick_start/pricing/

Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.

I’ve found Flash 0731 to be pretty great recently. I do feel like a lot of the models are pretty close in terms of ability. I often run code through multiple different models _and_ harnesses for code reviews and they all typically find the same things as each other.
Before DeepSeek-V4-Pro-0813's price goes up, I expect a surge of frantic traffic — hope the servers can hold up.
This model is not very good at coding, but it is quite good at research, evaluation and action, I don't write code, but it really goes head-to-head with the most expensive models in searches such as stock market and forex
The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?
Still behind Kimi-K3 in almost half of the benchmarks
DeepSeek is moving to peak/off-peak API pricing. Off-peak rates are 50% of peak rates.

Peak: 01:00–04:00 UTC and 06:00–10:00 UTC Off-peak: all other hours

New pricing takes effect August 16, 2026 at 16:00 UTC.

Model Period Cache hit Cache miss Output (input / 1M) (input / 1M) (/ 1M)

deepseek-v4-flash Off-peak $0.007 $0.22 $0.66

deepseek-v4-flash Peak $0.014 $0.44 $1.32

deepseek-v4-pro Off-peak $0.022 $0.66 $1.98

deepseek-v4-pro Peak $0.044 $1.32 $3.96

For batchable workloads, scheduling outside those two UTC windows cuts token costs in half.

I had high expectations for V4 Pro, especially since DeepSeek V4 Flash 0731 performed so well compared with other Flash models. What a letdown.
Been testing this on hobby project https://github.com/arj03/seedkernel/. Latest flash was a big step up. Pro feels really slow compared. Claude opus is still better day this level.
You mean to tell me it's been 10 hours and there's no unsloth quant? I've been running 0731 and love it.
Deepseek V4 Pro 0813 is the most unreliable model I have tried, it works on pass@3 shockingly well you can get it to match Sol or Fable perhaps in task done, but it's horrendous at pass@1 very prone to going wrong and doing horribly at most benches.

I am not sure what it is buy I suspect it might be GRPO.

Welp gonna give Deepseek more money. This is very cheap indeed. And I’ve been using them and kimi for a bit now not via open router but on my own and have found them on part with sonnet 5 though sonnet 5 these days I think has gotten worse.

At work I had to move to Fable to get decent work results.

Even though cost-per-token is low, Deepseek v4 tends to burn an immense number of tokens to accomplish tasks.
These people can't version control properly, V4.1 or V5 would be more appropriate.
Graphs without labels and/or scales on the axes are useless. I know less after viewing that page than before, but I got to see some pretty lines that I guess must mean something.
So flash is 52 points on artificial analysis, and pro is 53
OpenRouter is a place with zero support.

Suddenly get a big debt on your account with nobody to respond.

As an early adopter of OpenRouter, I'm afraid they are in shambles.

V4 Pro has vision correct?
Why do I feel that the Pro version's effects are inferior to Flash's? Is it just my imagination?
I'm Satisfied with this model (in opencode)
DS Pro is what I hoped it would be. I have used it as an auditor for a couple of implementations, and the work is solid. This will allow me to split the work between 5.6 Sol and DS Pro.
Did you guys read the fine print they plan to increase prices significantly in the future
Is having padded version numbers with a leading zero a common thing?

Wondering, sorry if it's a dumb triviality to ask.

Is this even a (sub-)version number? I mean the major version is clearly 4.

mannnn i just downloaded 0731 ffs
Why does this link to OpenRouter, which has no useful information on its own? Linking to the official API or the benchmarks would make more sense:

- https://api-docs.deepseek.com/

- https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)