▲ 8 points
back
2 comments
"We present SpQR, which allows lossless LLM inference at 4.75 bits with a 15% speedup. You can run a 33B LLM on a single 24GB GPU fully lossless. SpQR works by isolating sensitive weights with higher precision and roughly doubles improvements from GPTQ"
-- Tim Dettmers
https://twitter.com/Tim_Dettmers/status/1666076553665744896?...
code here: https://github.com/Vahe1994/SpQR