Speaker Adaptive Mixture of Weight-Decomposed LoRA Experts for On-Device End-to-End ASR
Qiuming Zhao, Guangzhi Sun, Chao Zhang, Mingxing Xu, Thomas Fang Zheng · IEEE Transactions on Audio Speech and Language Processing · 2025
On-device speaker adaptation faces significant challenges in balancing recognition performance with computational efficiency. To address these challenges, Mixture-of-Experts (MoE) is a promising approach, which enhances model representation power and expands model capacity through a dynamic gating mechanism, thereby better handling speaker variability. However, conventional MoE models are often very large, making them challenging to deploy on resource-constrained edge devices. In this paper, we propose a novel speaker adaptive mixture of weight-decomposed LoRA experts (SAMD) approach, which uses low-rank adaptation (LoRA) modules as experts to reduce the number of trainable parameters in MoE, and further decomposes the experts into magnitude and directional components to enhance both the learning capacity and training stability of the experts. Specifically, SAMD is applied to the quantised and personalised end-to-end automatic speech recognition models, which combines test-time speaker adaptation to improve the performance of heavily compressed models in speaker-specific scenarios. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size, 31.6% and 33.8% relative word error rate reductions were achieved on the quantised Whisper model and Conformer-based attention-based encoder-decoder ASR model respectively, comparing to the original full precision models.