Separated Noise Suppression and Speech Restoration: Lstm-Based Speech Enhancement in Two Stages
Maximilian Strake, Bruno Defraene, Kristoff Fluyt, Wouter Tirry, Tim Fingscheidt · 2019
Regression based on neural networks (NNs) has led to considerable advances in speech enhancement under non-stationary noise conditions. Nonetheless, speech distortions can be introduced when employing NNs trained to provide strong noise suppression. We propose to address this problem by first suppressing noise and subsequently restoring speech with specifically chosen NN topologies for each of these distinct tasks. A mask-estimating long short-term memory (LSTM) network is employed for noise suppression, while the speech restoration is performed by a fully convo-lutional encoder-decoder (CED) network, where we introduce temporal modeling capabilities by using a convolutional LSTM layer in the bottleneck. We show considerable performance gains over reference methods of up to 0.26 MOS points (PESQ) and the ability to significantly improve intelligibility in terms of STOI for low-SNR conditions.