Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
Published in Annual Conference on Machine Learning and Systems (MLSys), 2026
This paper introduces an algorithm-system co-design for accurate 2-bit KV cache quantization, significantly reducing GPU memory consumption and increasing inference throughput during LLM inference.