(Input / Output / Cache Read, [$/M])
DeepSeek-V4-Flash:
Prev: 0.14 / 0.28 / 0.0028
Off-Peak: 0.22 (1.6x) / 0.66 (2.4x) / 0.007 (2.5x)
Peak: 0.44 (3.1x) / 1.32 (4.7x) / 0.014 (5.0x)
DeepSeek-V4-Pro:
Prev: 0.435 / 0.87 / 0.003625
Off-Peak: 0.66 (1.5x) / 1.98 (2.3x) / 0.022 (6.1x)
Peak: 1.32 (3.0x) / 3.96 (4.6x) / 0.044 (12.1x)
gpt-5.6-luna: $0.20 / $1.20 / $0.02 / $0.25 (In / Out / Cache Read / Cache Write)EDIT: formatting
EDIT2: giving up on the formatting :-/
DeepSeek was hugely underpricing cache hit pricing before and even after this increase they're still cheaper on that metric than every other provider I'm aware of, but it will put an end to those "I used 1 billion tokens and spent $4" reports.
https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
If I'm reading the benchmarks right, they now went from being much cheaper than Luna (but twice as slow), to being roughly same price (but twice as slow).
So all else being equal, where I would previously have used DeepSeek, I can just use Luna, and get the same result twice as fast?
(Yeah I know benchmarks are mostly nonsense, but the ones measuring time are real, and it's the most precious resource.)
Right now, they give 4100 credits for Luna and 63 000 for Deepseek on their prepaid plan (both are 2x)
I pasted the same prompt into OpenCode, set to Deepseek v4 flash free and did it first try.
I'm was going to purchase Opencode GO to try it, but seems my timing is really bad :( hope it doesn't go up too much in Opencode or they find other providers. Bad timing!
I am a person that buys into a tool or a process and expects it to be part of the life with no major changes through the years (or as long as the need exists). But AI? You buy into something today, not 2 weeks have passed and there's already a large "update" introduced to the conditions or the optimal usage patterns you should be adopting.
It's tiring. Makes all prices and offers feel so unreliable and gets me a bit more disinterested each time they change.
The same motivation of course - the GPUs have a finite service lifetime, so to maximize revenue you need to keep them busy 24x7.
https://www.baseten.co/pricing/
If anyone has tried Baseten versions of these Chinese frontier models, let me know what you found.
DeepSeek-V4-Flash (off-peak, x2 for peak)
* Cache Hit $0.007 (x2.5)
* Cache Miss $0.22 (x1.5)
* Output $0.66 (x2.25)
DeepSeek-V4-Pro (off-peak, x2 for peak)
* Cache Hit $0.022 (x6)
* Cache Miss $0.66 (x1.5)
* Output $1.98 (x2.25)
Peak Hours: 01:00–04:00 and 06:00–10:00 UTC
Effective from: 16:00, August 16, 2026 (UTC)