Accelerating NLP with Token Pruning: A Survey of Methods and Applications
Wei Li, Yimin Wang, Hao Zhang, Chen Jie, Shui Xiuying, Fang Liu, Ming Yang, Lei Zhou, Sun Rui, Zhen Huang · 2025
Transformer-based models have revolutionized natural language processing (NLP) by achieving state-of-the-art performance across a wide range of tasks. However, their high computational cost remains a major challenge, particularly in real-time and resource-constrained environments. Token pruning has emerged as an effective technique for reducing inference complexity by selectively removing less important tokens during processing. This survey provides a comprehensive review of token pruning methods, categorizing them into heuristic-based, learnable, and reinforcement learning-based approaches. We discuss the theoretical foundations of token redundancy in transformers, analyze pruning strategies applied at different stages of model execution, and summarize empirical results demonstrating the trade-off between efficiency gains and accuracy retention. Additionally, we explore real-world deployment considerations, including hardware compatibility, task-specific pruning performance, and robustness challenges. We conclude by outlining open research directions, including adaptive pruning strategies, multimodal extensions, and theoretical advancements. By consolidating existing knowledge on token pruning, this survey aims to guide future research and practical implementations of efficient transformer-based NLP models.