Quantization?
back
1 comments
Its not a new model, but rather their infrastructure and hardware they are showcasing.
Groq appears to have quantized the Kimi K2 model they're serving, which is part of the reason why there's a noticeable performance gap between K2 on Moonshot's official API and the one served by Groq.
We don't know how/whether the Qwen3-235B served by Cerebras has been quantized.
Cerebras have previously stated for other models they hosted that they didn't quantise, unlike Groq.