back
2 comments
"We present SpQR, which allows lossless LLM inference at 4.75 bits with a 15% speedup. You can run a 33B LLM on a single 24GB GPU fully lossless. SpQR works by isolating sensitive weights with higher precision and roughly doubles improvements from GPTQ" -- Tim Dettmers

https://twitter.com/Tim_Dettmers/status/1666076553665744896?...