10M Tokens LLM Context(github.com)
github.com
10M Tokens LLM Context
https://github.com/SqueezeAILab/KVQuant
1 comments
KVQuant: Towards Enabling 10 Million Context Length For LLM Inference through KV Cache Quantization
https://github.com/SqueezeAILab/KVQuant