OS-Denseformer: A Lightweight End-to-End Noise-Robust Method for Chinese Speech Recognition
Shiqi Que, Liping Qian, Mingqing Li, Qian Wang · Applied Sciences · 2025
Automatic speech recognition (ASR) technology faces the dual challenges of model complexity and noise robustness when deployed on terminal devices (e.g., mobile devices, embedded systems). To meet the demand for lightweight and high-performance models in terminal devices, we propose a lightweight end-to-end speech recognition model, OS-Denseformer (Omni-Scale-Denseformer). The core of this model lies in its lightweight design and noise adaptability: multi-scale acoustic features are efficiently extracted through a multi-sampling structure to enhance noise robustness; the proposed OS-Conv module improves local feature extraction capability while significantly reducing the number of parameters, enhancing computational efficiency, and lowering model complexity; the proposed normalization function, ExpNorm, normalizes the model output, facilitating more accurate parameter optimization during model training. Finally, we employ distinct loss functions across different training stages, using Minimum Bayes Risk (MBR) joint optimization to determine the optimal weighting scheme that directly minimizes the character error rate (CER). Experimental results on public datasets such as AISHELL-1 demonstrate that, under a high-noise environment of −15 dB, the CER of the OS-Denseformer model is reduced by 9.95%, 7.97%, and 4.85% compared to the benchmark models Squeezeformer, Conformer, and Zipformer, respectively. Additionally, the model parameter count is reduced by 53.35%, 10.27%, and 27.66%, while the giga floating-point operations per second (GFLOPs) are decreased by 67.51%, 66.51%, and 13.82%, respectively. Deployment on resource-constrained mobile devices demonstrates that, compared to Conformer, OS-Denseformer reduced memory usage by 10.79% and decreased inference latency by 61.62%.