back
3 comments
If you look at the demo video, the output is exactly the same for both cases, so I doubt it uses quantization.
Exactly what I was thinking. Everyone already does this. Unless they’re doing something else, they’ll have to show why it’s better than just quickly quantizing to 8 bits or 4 bits or whatever.
Whatever it is, it will likely be copied into the open source tooling like llama.cop soonish or something similar will arrive in llama.cpp. It doesn’t seem defensive advantage. It seems like a feature and fighting against fast moving open source alternatives.