Bone and Air-Conducted Speech Fusion Based on LSTM Network and Kalman Filtering

Zhenglong Liu, Zhe Chen, Dapeng Yu, Fuliang Yin · IEEE Sensors Journal · 2025

Integrating bone-conducted (BC) and air-conducted (AC) microphones can synergistically enhance noise suppression and improve speech quality. For this purpose, a novel time-domain BC and AC speech fusion enhancement method based on the long short-term memory (LSTM) network and Kalman filtering is proposed. Specifically, the line spectral frequencies (LSFs) and the corresponding residual power of the clean AC speech are predicted using an LSTM neural network (NN) from noisy BC and AC speech parameters. Then, the joint state and observation models about BC and AC speeches are established with the estimated parameters through the linear-predicted-based speech model. Finally, the AC speech is enhanced through Kalman filtering by fusing noisy AC and BC observations. Simulation results on the elevoc simultaneously recorded microphone/bone (ESMB) dataset and self-recorded dataset illustrate that the proposed method obtains good speech enhancement (SE) performance and generalization ability with lower computational complexity compared with other existing methods, improving the speech quality by 1 point in the perceptual evaluation of speech quality (PESQ) and 0.1 point in short-time objective intelligibility (STOI) under the 5-dB white noise condition. Real-world experiments further confirm the effectiveness of the proposed method.

Read the paper · More papers on PaperTik