I get over 100 tok/s sustained on my M4 Max and M5 Max, in MacBook Pro's. LM Studio + MLX.
back
2 comments
Same experience on M4 Max .. but quality of qwen still leaves so much to be desired after getting used to virtually unlimited tokens at work. Many people on this (and similar) thread seem to believe local models would inevitably improve, and I want to believe this too, but I don’t see this ever happening without growing in size
With Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf?
Also, funny lumping the M4 "and" the M5, I find them 15% to 45% different performance, depending.
And for a good deal of work, an M3 Studio Ultra outpaces the M4 and ties the M5 on single work at a time, outpaces both doing multiple work at a time.
> With Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf?
Ahh no I'm using the MLX version, it's about 5-10% faster than GGUFs in my experience.