They claim average 1.45x speedup and maximum 1.7x speedup over GPTQ.
High-level idea: About 1% of the weights contribute greatly to quantization error. So skip the quantization of these weights.
High-level idea: About 1% of the weights contribute greatly to quantization error. So skip the quantization of these weights.