"Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning" by Heng Wang , Jielin Qiu , Wenting Zhao , Cheng Qian , Liangwei Yang , Jiawei Han , Heng Ji , Silvio Savarese , Shelby Heinecke , Huan Wang
TLDR:
The text discusses how random eviction of reasoning tokens without scoring matches selective KV cache compression due to the self-protecting nature of reasoning traces through redundancy, eliminating the need for scoring once prompts are preserved. Inkling-Small Large language models excel at tasks requiring extended reasoning but face memory bottlenecks with long thought chains in the KV cache. Unlike existing methods that score tokens based on future importance, the proposed Random Attention approach evicts uniformly at random, achieving higher throughput and matching prior methods in deployment. The selection signal is found to be minimally impactful, as redundancy in the reasoning trace within the text and across attention heads ensures the prompt's safety and retains necessary information without the need for scoring. The study's findings and code are publicly available for reference.
Read Paper / Blog

huggingface.co
Paper page - Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
Join the discussion on this paper page
1September 4, 2026 71 2