Band-Wise Front-End Distortion Suppression for Robust Speech Recognition
Siyi Zhao, Wei Wang, Yanmin Qian · 2024
Advancements in deep learning techniques have significantly improved automatic speech recognition (ASR). However, improving robustness to acoustic interferences, such as background noise and reverberation, remains crucial for accurate speech recognition in noisy environments. Traditional speech enhancement (SE) modules used as pre-processing front-ends often fail to integrate well with ASR models due to intelligibility distortions introduced by the SE module. Inspired by the observation adding technique and band-wise network design of recent SE models, we propose a novel adaptive band-wise observation adding (BOA) module for front-end distortion suppression. The lightweight BOA network learns time- and frequencyvariant OA factors, adaptively balancing the contributions of noisy and enhanced speech signals for diverse noisy conditions. Experiments conducted on both simulated and real-world test sets demonstrate that the BOA module produces more effective dynamic OA factors under various conditions compared to fixed OA factors, leading to improved ASR performance.