OpenCode Go even has double limits temporarily so for 10 USD you effectively get 140 USD of tokens to spend. It would impress me if someone could burn that amount with "normal" usage. Even when running multiple sessions.
I have a Claude Max subscription but I've barely touched it, it just feels like a step back to have to think about limits and usage even though the models are stronger.
The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues. Test coverage too low? Auto generate tests on CI for every pull-requests! Monitoring server logs, continuous security audits and investigating every received exception now becomes possible.
I'm thinking about having it automatically filter and re-rank my social media feeds so I can steer the algorithm instead of the other way around.
Perhaps other people (with enormous budgets) were already doing all of the above but for us this is a really exciting release!
There's no way large companies outside the US will pay the "US AI lab" premium if they can get the same workloads done at a fraction of the cost using open-weight models that they can self-host and optimize/fine-tune on.
Another time, Flash started trying to make tool calls by just calling bash and catting the tool call to stdout. Then it started running echo xx for every two letter UNIX command it could think of: mv, cp, etc and the it dug into uv, ty, and jj
which in your case is?
And, probably 99.99% of people using LLM probably don't even need SOTA anyway.
The 'floor' has gone up: today's model a bit behind SOTA is like model releases that were blowing people's minds a few months ago. Compared to, say, DS R1, this is far out stuff.
This, Luna, and (if it's good in practice) Laguna S are also fast and light not just cheap. And, as happened before, DeepSeek's first but other open model makers likely follow.
And a small, fast model taking small steps is...fun? More like working with code.
My initial thought was to sign up for ChatGPT, but I had $20 in OpenRouter so I've been trying out DeepSeek V4 Pro with Pi for the last few days and I gotta say, it's good enough for my use case. And even with paying for API usage rather than Claude's subsidised subscription, and with OpenRouter taking their cut, I will probably end up paying significantly less overall. And I really like the flexibility of being able to use whatever minimalist open source harness I want (and being able to switch providers easily, too).
(My demands probably aren't as high as many others' - I mostly use it for help with some hobbyist coding projects, and I tend to ask it questions about how to approach problems rather than just telling it to go off and code stuff for me.)
I've been running this model locally for a week, and the preview version before that. This updated one feels like a whole tier up. It's very capable for debugging and analyzing documents/data I upload.
The killer feature, IMO, is the speed. On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s on a single stream. I saw 1000 tok/s with ~64 concurrent streams on vLLM.
That's fast enough that you can interactively chat with it without switching tabs while you wait, and its a ~300B (13B active, hence the speed) model so the responses are also very good. It's actually more convenient now for me to direct 95%+ of my day to day usage to my local model, and only use Claude Fable for really big coding tasks.
Until this model was released, I was contemplating spending even more money on hardware to run GLM5.2 (~750B) at reasonable speeds, but I no longer feel that need. This is smart enough, and I think it only gets much better for local models from here.
It is strong (not Fable strong though) with a much better “persona” than Opus, and very different blindspots. If you flip between Claude and this you will find both catch the mistakes of the other before they get out of control.
On balance I actually prefer DeepSeek for programming now, because of the way it talks.
This is on Pi agent, nothing fancy at all about my prompts or use case. Anyone else experiencing this?
I've also had it randomly go from talking about Rust to talking about the electric chair, controversies about D&D rules (both irrelevant and something I've never discussed) and it's completely blind to it in future prompts even when its pointed out and referenced directly
All this said its still worth it but the agentic performance has degraded in my experience at least
Which would put them... exactly where everyone else is on this graph.
Edit: I seem to have misunderstood the news. I thought the magical cache read pricing was going away (0.002) and they were going to be on par with everyone else (0.02). But I have no idea.
Edit 2: Apparently, neither do they!
>We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
And how this has been accelerating!!
I felt this very hard when I had to travel in the middle of nowhere in south america, with no network, and wanted to keep an LLM model on my macbook pro with 48GB of RAM. That was back in April 2026, a few months ago.
I downloaded Google Gemma 4 (google/gemma-4-26b-a4b) and - Oh boy - I was amazed by it's capacity!
I was able to use it to code simple things, ask it about nature, learn new stuff while traveling and make stories for the kids.
Was really amazing to observe and experiment this!
Seems to me there will be some good chance to run these great LLM locally on our hardware!
Amazing time to be alive
I Compared Deepseek V4 Flash 0731 (low) to Gemini 3.5 Flash Lite (minimal) and GPT 5.6 Luna (no reasoning) and Deepseek V4 Flash 0731 gets it wrong alot, where as Gemini and 5.6 Luna just gets it done.
Someone else here said we could get the model via OpenCode Go for $10/mo and get about $120 worth of credit, so I decided to give it a whirl. My first month is actually $5.
In the 3 hours I've been using it I've burned 3% of my 5-hr, 1% of my weekly and 0% of my monthly.
It's fixed 3 or 4 issues in my C++ game, despite not having visual capabilities to see the screenshots I was trying to give it. One-shotted them too.
Luna struggled with what I thought was an easy task (had to replace a few ASCII chars with the correct unicode char but kept choosing incorrectly).
I'll keep using it.
I have £20/month Gemini and £20 a month claude for a bunch of personal projects.
Yes I have to wait sometimes, it's probably a good thing.
Furthermore, in my company we are using MCPs for Google Ads (it manages our ads), Analytics, Search Console, Zoho CRM, Microsoft Clarity... We use it to crawl specific websites and send daily summaries to our sales team in MS Teams channel. We use it to send daily summaries on marketing statistics and analytics... All with a FREE model. We are rarely hitting any limits so far and in case we need more tokens - we use NOUS or openrouter to pick between Flash or Pro for specific tasks that require more churning.
AMA.
From here on, it's going to become all about harnesses that best situate and organize swarm intelligence at scale.
[0] https://taylor.town/silver-landmines
When I see dramatic leaps like this, it tells me that the important hacks haven't yet been discovered.
But note that you have to use Cline (or other harness) if using vscode. I was shocked at how poor the recent versions of GitHub Copilot are at using the cache (with Fireworks AI, but I believe it's a more generic problem).
And a meaningful chunk of the comments are saying "this piece of garbage isn’t even at the level of gpt-oss 20B".
If I'm reading the chart correctly, a couple observations:
* deepseek-v4-flash-0731 max is better than kimi-k3 max
* glm-5.2 is dumber than a box of rocks (this must be on low reasoning or something, right?)
This is way more extreme than other results I'm seeing, like those from Artificial Analysis.
Does no thinking emissions for context saving.
Btw if you need an app I may deliver it to you in ten minutes for just five cents if I'm in the mood. Just let me know.
ARC-AGI II:
- GPT-5.2 (medium) %26.7 ($0.759)
- DSV4-Flash (max) %61.4 ($0.04)
Not a huge deal since it's still cents per session, but my bigger issue was the weird change in tone. It became a lot more pretentious and over-explanatory.
Heavy prompt reworking helped but maybe that's just the cost of being better at coding and ARC-AGI?
When I need vision capabilities I use GPT 5.3 codex and if deepseek can’t figure something out after a few goes I switch to GTP 5.5 or 5.6 (I’ve been giving Terra first bite recently and it does pretty well, and have used Sol a couple of times).
Using this regimen means I spend under $100 per month on inference and I work all day everyday with multiple agents running simultaneously all on API token spend not subscriptions.
What secret sauce do they have?
But it makes me quite curious, how a text-only model can do so well on ARC-AGI-2 being a set of visual puzzles? It would have to solve it entirely using text-only spatial reasoning about the grid (or maybe writing code?). I am curious if this is normal or do other models use their vision capabilities to solve the puzzles?