Efficient Transformer Models via Language-Aware Frequency-Based Vocabulary Pruning

Mukhlish Fuadi, Adhi Dharma Wibawa, Surya Sumpeno · IEEE Access · 2026

Multilingual Transformer models offer effective cross-lingual generalization capabilities. However, their architecture suffers from embedding-parameter overhead due to massive vocabulary sizes, where most embedding parameters are redundant for specific monolingual tasks. This study investigates language-aware frequency-based vocabulary pruning as a deterministic and reproducible strategy to mitigate this inefficiency without architectural modifications or additional large-scale pre-training. We evaluate the mDeBERTa-v3-base model under monolingual (Indonesian) and hybrid (Indonesian–English) vocabulary configurations on Named Entity Recognition (NER) and Sentiment Analysis (SA) tasks. Multi-seed experiments show that reducing the vocabulary to 30k tokens lowers peak GPU memory usage during inference from 6.64 GB to 4.70 GB (≈29%), while maintaining comparable inference latency and minimal performance degradation, with an average drop of less than 1.5% across entity-level (NER) and sentence-level classification (SA) F1-scores. Further analysis attributes the hybrid model’s stability to substantial cross-lingual subword overlap (over 24,000 tokens) and efficient tokenization behavior, reflected in a subword fertility of approximately 1.008, indicating minimal word fragmentation. Statistical significance testing confirms that the hybrid model’s performance on sentence-level classification tasks shows no statistically significant difference compared to the baseline (p> 0.05). Although evaluated on Indonesian as a representative case study, the results illustrate the potential of this language-agnostic framework as a practical solution for optimizing Transformer models in resource-constrained deployment scenarios.

Read the paper · More papers on PaperTik