A Front-End Adaptation Network for Improving Speech Recognition Performance in Packet Loss and Noisy Environments

Yehoshua Dissen, Shiry Yonash, Israel Cohen, Joseph Keshet · IEEE Transactions on Audio Speech and Language Processing · 2025

Robust automatic speech recognition (ASR) in packet loss and noisy environments remains a significant challenge. Large pretrained transformer models have made notable strides in improving ASR performance across diverse domains. However, considerable room remains for improvement, even in moderate packet loss and noise conditions. Enhancing these models is particularly difficult because retraining is computationally prohibitive, and fine-tuning introduces the risk of domain shift, which can degrade performance in other languages or environments. We introduce a novel method that leverages a front-end adaptation network to improve word error rate (WER) performance in scenarios with packet loss and noise. Our approach addresses the constraints of working with large pretrained ASR models while avoiding retraining or fine-tuning. We connect an adaptation network to a frozen ASR model, where the network is trained to modify corrupted input spectra using both the loss function of the ASR model and an enhancement loss. This strategy allows the system to adapt to packet loss and noise without compromising the performance of the original ASR model or generalization across domains. The method focuses on improving WER rather than signal quality or intelligibility, targeting it for ASR applications. We conduct a comprehensive set of experiments on various types of noise. Our results demonstrate that the adaptation network significantly reduces WER in all conditions while preserving the foundational performance of the pretrained ASR model.

Read the paper · More papers on PaperTik