Microphone array speech enhancement using LSTM neural network

Anton Buday, Jozef Juhár, A. Cizmar · 2019

The article encompasses microphone array speech processing using neural networks. Noisy microphone array, which consists of 12 elements, is simulated from clean and noise mono-channel speech recordings with the utilization of open custom-modified software framework MCRoomSim, which is executable in an integrated development environment called MATLAB. The modified framework applies beamforming methods, e.g. Frost algorithm in order to suppress noise signal, this is known as primary speech enhancement. Such beamformed signal is filtrated by the application of the Wiener filter, which is predicted from noisy speech spectrograms using a deep neural network model. This neural network predicted Wiener filter, originally calculated out of spectrograms, is subsequently multiplied with beamformed signal for the purpose of secondary speech enhancement. The latter way of speech enhancement is generally called beamforming with post-filtering. There are various parameters for objective evaluation of speech enhancement effectivity, whether concerning the very beamforming or application of neural network, i.e. STOI, PESQ, MOS-LQO, and even fwSNRseg. The beneficiality of beamforming is discussed in the last chapter of this paper.

Read the paper · More papers on PaperTik