OScaR-KV-Quant. 🏆 OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond — redefining the accuracy-efficiency Pareto front for X-LLMs KV quantization.
138Awesome-Attention-Sink. 🚀 First survey on Attention Sink in Transformers — 200+ papers on utilization, interpretation, and mitigation.
137Super-Experts-Profilling. (ICLR 2026) Unveiling Super Experts in Mixture-of-Experts Large Language Models
43RotateKV. Code used to reproduce the simulation results of RotateKV.
5Awesome-Vision-Token-Pruning-for-VLMs.
4Awesome-Diffusion-Quantization.
1