What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.
When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left
And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome
Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.
I have screenshots of both. The description above the chart is the same in boh cases:
> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)
What happened? How can the scores change so much in a few seconds?
https://artificialanalysis.ai/methodology/intelligence-bench...
Edit to provide AA's article explaining it:
https://artificialanalysis.ai/articles/artificial-analysis-i...
The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.
Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-i...
I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.
Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol
Source: http://pellmell.ai/leaderboard.
This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.
I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc.
It is smart and reasonably quick but not reliable.
I wasn't able to find an explanation from them. Anyone knows what happened?
Qwen3.6-27B-FP8 currently espouses (see below), which is not too bad. The directions are a bit mad but the mileage is about right and there is a helicopter manufacturer here and a RNAS (navy not airforce) museum nearby at Yeovilton. Cosford is in Shropshire which is not a million miles away.
I'm not sure what 盆地的 means but the river Yeo is correct ... OK ... "basin like" - again not bad, even if Chinese is not the first language here. The model understands that Yeovil is named after (or vice versa or at least is associated with) a river
Yeovil is the current form of Gifle (Saxon) which I thought meant "bend in a river" but WP is currently saying "fork in a river". My source is a local museum. There is a fork but was it there 2000 odd years ago? My hydrology skills say ... possibly
---------------------------------------------------- Q: where is yeovil:
Yeovil is a town in Somerset, in the South West of England.
It is located roughly:
25 miles (40 km) south-west of Exeter
60 miles (100 km) west of Bristol
140 miles (225 km) west-south-west of London
Yeovil is known for its historic market town center, RAF Museum Cosford (nearby), and as a significant industrial town, particularly during World War II for aircraft manufacturing (including the Wellington bomber). It sits in the盆地的 valley of the River Yeo.Clearly the weighting of those things depends on the usecase