License pretty similar to k3 with some caveats. Free to use for internal or <50M$ revenue / year. Limitations above that threshold for serving the model or services targeting coding / productivity agents.
Benchmarks are looking good, trading blows w/ opus4.8 and sol, generally 10-20p under fable. But that's neither here nor there w/ qwen, their benchmark to real world usage correlation has been iffy in the past.
The local model 3.8-27B announced for Friday, same time so ~48 hours from now. That'll be a bit more exciting for a lot more people, since 3.6 was quite good for local inference, and their 3.7-max -> 3.8-max shows a lot of improvement.
It uses twice as much tokens to achieve the same but the results are significantly better and because it's so much cheaper it's the most economical choice too.
[1] https://blog.bosun.ai/software-maintenance-with-open-weight-...
Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:
- xhigh (default): for complex tasks demanding thorough analysis
- medium: balancing accuracy and speed
- low: efficient reasoning optimizing for speed and cost
In addition, preserve_thinking is enabled by default for all workloads for the best out-of-the-box experience.
Asking because in my case (OCR of scanned historical "National Geographic" magazines) the LLM trying to merge text split into separate columns was running in circles from time to time and needed a lot of prompt tuning when using Qwen 3.0/3.5/3.6 (still needs from time to time).Maybe I’m misreading this or some other post, I thought QWEN was stepping away from releasing these models for local consumption
[1] https://www.reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepsee...
In terms of what you get for what you pay for, it's incredible - probably by far the best.
But unless I'm reading things wrong, it does not appear to be top-of-the-line.
The 1bit quant model is at an astonishing 397GB with 95B active per MOE. This literally puts Opus 4.5 performance level into a machine a normal person could buy, and still gets usable tokens/second.
The full lossless model BF16 is clocking at 4.9TB. The model card claims the model to be between Opus 4.8 and Fable 5. Again that's astonishing as getting a machine with 7TB RAM (with context + KV cache) is still within the realm of medium size companies.
Bad things: The open source version has its vision capability removed, and the context capped at 250k . I expect someone to bolt a Kimi 2.6 vision tower to it to restore the vision capability (at less performance of course). For context, I played around with extending the context to 600k for Qwen 3.5 397b, and the context remained stable up to around 480k. It'd be interesting to see if the same can be done to Q3.8 .
Also no out of the box DSpark/DFlash support. MTP is present so we should at least get some boost in TP speed.
Honestly this model people at home can tinker with, if you have a big enough Mac. Maybe 4 Strix Halo/DGX Spark, and then at 1 bit quant? Nah.
Use the right sized model, for your hardware. You'll get better results.
Suppose I have 100GB of unified memory, how should I know which model suits it best? I understand how a 2.4T model wouldn't fit, but I don't understand the impact of quantization and whether I should use a 200G model quantised to fit say 90GB of memory, or a non-quantised 90G model.
at this kind of quantization is it useful though?
That is unfortunate, that the open weight model doesn't have vision support or the 1M context length...
[1] https://old.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_...
The whole industry is now pushing through memory.
In 5 years you have either some type of explosion which willjust make all the hardware from today affordable or you have such an AI explosion, that the today hardware is written off and not efficient enough anymore that you can buy it for cheap.
In parallel, its clear that we need more memory.
In parallel models in hardware will become a thing on mass market.
In parallel everything gets more efficient. The 30B parameter model will be for sure more intelligent in 5 years than it is today.
Read the room, Qwen. It's not a good time to hobble your releases.
Make of this what you will.
I'm interested in your take on it. IIRC Gemma family models too have a ~250k vocabulary size
[0]: https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-se...
[1]: https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-se...
best crypto-bro impression I can do...