back
354 comments
China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you.

What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing.

When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left

And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.
Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.
I'm still skeptical of the smaller models after the talent exodus a few months ago.
Who is going to break it to the Americans that China is more than a slight favorite to win an existential battle over which country is better at math?
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

I have screenshots of both. The description above the chart is the same in boh cases:

> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

What happened? How can the scores change so much in a few seconds?

Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date.

The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.

Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-i...

They should probably freeze the results before publishing.
Welp. That didn't last long
Same, they just updated it. Hacker news effect?
I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth.

I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.

Did you also try Opus 5 and 5.6 Sol?
How CLI are you guys using for qwen and kimi?
Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
Strange that the page https://artificialanalysis.ai/agents/coding-agents doesn't even mention "Qwen" once if it's now the "best" according to one of their one index?
Opus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence

Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol

Source: http://pellmell.ai/leaderboard.

This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge people.. wayyyy overpriced.
Hopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.
I find that surprising.

I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.

It is smart and reasonably quick but not reliable.

A couple days ago they had published an overall score of 53 for this model, but that was removed and today it returned with a score of 56.

I wasn't able to find an explanation from them. Anyone knows what happened?

Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.
The fact that the Chinese models have caught up on benchmarks suggests to me that its likely we will start to transition now into much more of a brand war. It will be subjective qualities that drive our decisions more than measures of absolute intelligence. Already I am choosing models more because I like the personality or style of what they do than because I think they have the absolute highest chance of outputting the most technically correct answer to any given prompt. It will be very interesting to see how things evolve in this direction.
My vague equivalent of the pelican riding a bicycle test (for a local model without internets) is to ask it: "Where is Yeovil"? I don't expect a totally accurate answer for obvious reasons but I do enjoy watching the accuracy improve.

Qwen3.6-27B-FP8 currently espouses (see below), which is not too bad. The directions are a bit mad but the mileage is about right and there is a helicopter manufacturer here and a RNAS (navy not airforce) museum nearby at Yeovilton. Cosford is in Shropshire which is not a million miles away.

I'm not sure what 盆地的 means but the river Yeo is correct ... OK ... "basin like" - again not bad, even if Chinese is not the first language here. The model understands that Yeovil is named after (or vice versa or at least is associated with) a river

Yeovil is the current form of Gifle (Saxon) which I thought meant "bend in a river" but WP is currently saying "fork in a river". My source is a local museum. There is a fork but was it there 2000 odd years ago? My hydrology skills say ... possibly

---------------------------------------------------- Q: where is yeovil:

Yeovil is a town in Somerset, in the South West of England.

It is located roughly:

    25 miles (40 km) south-west of Exeter
    60 miles (100 km) west of Bristol
    140 miles (225 km) west-south-west of London
Yeovil is known for its historic market town center, RAF Museum Cosford (nearby), and as a significant industrial town, particularly during World War II for aircraft manufacturing (including the Wellington bomber). It sits in the盆地的 valley of the River Yeo.
This has me hopeful for Qwen3.8-27B!
Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
I distrust any benchmark where Opus 5 beats Fable 5.
Haven't tried this yet, but going to soon! I have to wonder what happened at Anthropic. We've cancelled our subscription in favor of OpenCode & Codex. Sol is just so good & OC goes so far for every $ spent. Claude's become a pain to work with - average output with an annoying personality. Who knew this would be an issue even a year ago? In any case, loving the stuff from the Chinese models!
Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?
It is also the most expensive open source frontier model, per task; cf. Cost per Intelligence Index Task. If it is as good as the benchmarks indicate it bodes well for Qwen and China. For my part, I'll pass; it is not on the Pareto frontier.
That took what 2 years? I love how the chip ban made them more efficient
It's the first model I've used that makes me forget it isn't one of the big frontier players after the first prompt. So far I'm impressed.
I just don't think you can combine speed, latency, price and intelligence into a single useful metric.

Clearly the weighting of those things depends on the usecase

Well, that's the bad index then. It is barely usable in my opinion compared to other Chinese frontier models.
Qwen is just crushing it overall. I regularly use 3.7-flash for everyday coding needs and it gets the job done.
It doesn’t even show up in the raw intelligence index, so how could it possibly be the best?
It's censored and it'll spread certain kind of narrative all over the world.
Out of curiosity, what's currently the best model I can use locally?
Anthropic have a real fight on their hands now. The competition is no longer 6 months behind, it's 6 days. If this had come out two or three weeks earlier this would be an absolute market leader on both quality and timeline.
There are great visual comparisons on this page.
curious about methodology. ive seen them post results for claude/codex when they only ran over benchmarks 3 times per model...
I'm seeing opus 59.2, Qwen 58.4?
Go China!
Qwen is the way to go
Does it matter, it's all non deterministic bs ware and deepseek is eating the Americans lunch