Fully Convolutional Neural Network-Based Speech Enhancement for In-Vehicle Environment in 3D Perspective
Kaikun Pei, Lijun Zhang, Dejian Meng, Wei Tian, Zhuang Zhang, Jianfeng Wu · 2025
Voice interaction is one of the important development directions of intelligent cabins. However, various noise interferences inside and outside the vehicle pose significant challenges to human-vehicle interaction. Speech enhancement technology can extract clear speech signals from mixed speech signals, thereby significantly improving the quality and intelligibility of voice commands, making it a current research hotspot. To address the issue of low intelligibility of voice commands in vehicle environment, we propose a speech enhancement method based on 3D tensor representation. Specifically, through short-time Fourier transform, the real and imaginary parts are concatenated in a new dimension to form a 3D tensor as the input feature. Based on convolutional neural networks, we conducted systematic research, building UNet speech enhancement models based on 1D, 2D, and 3D convolutions respectively, and compared the performance of the proposed 3D complex domain model with that of the time-domain and frequency-domain models. Additionally, we incorporated an attention mechanism and developed a time-frequency domain joint model by leveraging its information filtering and focusing capabilities, thereby significantly enhancing the speech enhancement effect. Finally, we conducted extensive experiments on the Voicebank+Demand dataset. The results show that although the time-frequency domain joint model outperforms the 3D complex domain model on the test set, in the vehicle environment, the 3D complex domain model achieved the highest scores of 3.67, 97.94, and 4.36 respectively in the three metrics of PESQ, STOI, and COVL.