Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
Published in Annual Conference on Machine Learning and Systems (MLSys), 2026
Algorithm-system co-design for accurate 2-bit KV cache quantization, reducing GPU memory consumption and increasing inference throughput during LLM inference. Modeled and optimized for efficient large language model serving.