back
1 comments
There's a tiny bit of play, like sliding window attention. As tokens leave the sliding window you can keep them or discard them. If you keep them, you can freely truncate the context and resume generation from an earlier point. If you discard them, you have to recompute the KV cache up to that point.

Llama.cpp checkpoints and moves snapshots of the cache to main RAM for faster resumption after truncation.