Flexible Language Models: A Comprehensive Review on Dynamic Pruning, Sparse Activation, and Modular Architectures for Sustainable NLP

Anjali Kshatriya, Komal D Prajapati · 2025

Large Language Models (LLMs) such as GPT, BERT, and T5 have transformed Natural Language Processing (NLP) and achieved state-of-the-art outcomes across domains including healthcare, education, legal analysis, customer service, and more. However, their extreme computational expense, inefficiency, and inflexible architecture present barriers to sustainable deployments, particularly in edge and low-resource environments. This review describes a number of recent advancements in efficiency-driven NLP research with a focus on techniques such as Parameter-Efficient Fine-Tuning (PEFT), Sparse Activation, Dynamic Pruning, and Modular Architectures. We synthesize and critically review eight key papers representing the years 2023-2025 to identify problems and opportunities such as in cases of thematic coding, poetry generation, evaluation modeling, machine translation, uncertainty quantification, scientific collaboration, and generation of visualizations. Following analysis and comparisons across studies to identify gaps, we present the Flexible Language Model (FlexLM/FlexiCore)-a unified modular framework incorporating sparse activation, dynamic pruning, hybrid attention, and domain-aware fine-tuning. Our data suggests that FlexLM may maximize computational savings up to 30-50% while enabling adaptable low-resource and uncertainty quantification capabilities. This review in turn considers advances in sustainable, adaptive, and accessible NLP to help researchers moving forward.

Read the paper · More papers on PaperTik