Heterogeneous Network Framework with Attention Mechanism of Speech Enhancement for Car Intelligent Cockpit Speech Recognition
YingWei Tan, Xuefeng Ding · 2023
The success of deep learning has significantly benefited single-channel speech enhancement in terms of intelligibility and perceptual quality. Traditional approaches have primarily relied on a single model to predict the clean version of the speech signal. Considering the advantages of different model structures, we design heterogeneous network frameworks with attention mechanism. We operate at the waveform level. The weighted outputs of different systems are combined to acquire the final waveform. We propose two strategies, scalar weights and vector-based attention weights, to lead the allocation of weights, respectively. Additionally, we train the proposed model end-to-end. It is optimized in both the time and frequency domains, employing multiple loss functions to achieve the desired performance. Experiments are conducted on synthesized dataset in car intelligent cockpit environments. In terms of perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI) and scale-invariant source-to-noise ratio (SI-SNR), the results show the proposed framework achieves 1.87, 8.38%, and 18.43 improvements over the man-made noisy data in the speech enhancement experiment. Besides, the presented algorithms achieves 3.20% word error rate (WER) improvements over the same data in the speech recognition experiment.