Effort on this one?
I keep my agents on tight leashes, using them very interactively for bouncing off ideas, architecture, and then writing code (especially prototyping) and Flash has been crushing everything I ever needed it to do.
Maybe my ambitions are too tame compared to people needing Fable / Sol grade models, but I'm probably staying on Flash and not moving on to Pro for the foreseeable future.
Tested this model, and gpt-5.6-terra-high.
Results: this one had few issues. terra: none.
These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different.
They are good till the project is simple... not anymore.
I can't wait to try this new one.
I'm happy
Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.
Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.
| Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 |
| | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) |
|--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
| HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0 |
| Terminal Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 | 81.0 | 88.3 | 85.0 | 88.0 |
| NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 | 48.9 | - | 69.7 | - |
| Cybergym | 83.3 | 76.7 | 52.7 | 38.7 | - | 80.0 | 78.3 | 83.1 |
| DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 | 46.2 | 67.5 | 58.0 | 70.0 |
| Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 | 59.9 | 76.5 | 76.2 | 77.9 |
| Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 | 23.8 | 27.6 | 25.7 | - |
| AutomationBench (Public) | 31.8 | 25.1 | 12.8 | 10.8 | 12.9 | 30.8 | 27.2 | 29.1 |
| DSBench-FullStack | 71.1 | 68.7 | 41.8 | 37.0 | 61.8 | 73.7 | 71.6 | 77.2 |
| DSBench-Hard | 67.2 | 59.6 | 31.1 | 25.8 | 54.5 | 63.0 | 71.7 | 68.3 |
Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
Peak: 01:00–04:00 UTC and 06:00–10:00 UTC Off-peak: all other hours
New pricing takes effect August 16, 2026 at 16:00 UTC.
Model Period Cache hit Cache miss Output (input / 1M) (input / 1M) (/ 1M)
deepseek-v4-flash Off-peak $0.007 $0.22 $0.66
deepseek-v4-flash Peak $0.014 $0.44 $1.32
deepseek-v4-pro Off-peak $0.022 $0.66 $1.98
deepseek-v4-pro Peak $0.044 $1.32 $3.96
For batchable workloads, scheduling outside those two UTC windows cuts token costs in half.
I am not sure what it is buy I suspect it might be GRPO.
At work I had to move to Fable to get decent work results.
Suddenly get a big debt on your account with nobody to respond.
As an early adopter of OpenRouter, I'm afraid they are in shambles.
Wondering, sorry if it's a dumb triviality to ask.
Is this even a (sub-)version number? I mean the major version is clearly 4.
- https://api-docs.deepseek.com/
- https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)