Databases · Computer Science
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee +2
2025-10-01
Computation and Language · Computer Science
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
Zeyu Xing, Xing Li, Hui-Ling Zhen, Mingxuan Yuan +1
2026-01-29
Distributed, Parallel, and Cluster Computing · Computer Science
KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
Bo Jiang, Taolue Yang, Youyuan Liu, Chengming Zhang +2
2025-09-03
Machine Learning · Computer Science
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
Dmitry Akulov, Mohamed Sana, Antonio De Domenico, Tareq Si Salem +2
2025-09-22
Machine Learning · Computer Science
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
Xing Li, Zeyu Xing, Yiming Li, Linping Qu +5
2025-11-21
Computation and Language · Computer Science
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
Shiyu Ji, Yixuan Wang, Yijun Liu, Qingfu Zhu +1
2026-05-14
Computation and Language · Computer Science
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
Yixuan Wang, Qingyu Shi, Jiayu Zhou, Dianbo Liu +2
2026-03-18
Machine Learning · Computer Science
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney +3
2025-05-30
Computation and Language · Computer Science
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
Yuan Feng, Junlin Lv, Haoyu Guo, Yukun Cao +2
2026-05-29
Computation and Language · Computer Science
BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference
Junqi Zhao, Zhijin Fang, Shu Li, Shaohui Yang +1
2024-10-31
Computation and Language · Computer Science
ThinK: Thinner Key Cache by Query-Driven Pruning
Yuhui Xu, Zhanming Jie, Hanze Dong, Lei Wang +5
2025-02-28
Distributed, Parallel, and Cluster Computing · Computer Science
KV Cache Compression for Inference Efficiency in LLMs: A Review
Yanyu Liu, Jingying Fu, Sixiang Liu, Yitian Zou +3
2025-08-11
Computation and Language · Computer Science
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
Jiebin Zhang, Dawei Zhu, Yifan Song, Wenhao Wu +5
2025-02-21
Computation and Language · Computer Science
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
Youhui Zuo, Sibo Wei, Chen Zhang, Zhuorui Liu +2
2025-03-28
Machine Learning · Computer Science
KV Cache Offloading for Context-Intensive Tasks
Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy +1
2026-05-18
Computation and Language · Computer Science
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong +4
2024-07-26