Speech Separation in the Frequency Domain with Autoencoder
Hao D., Son Quang Tran, Duc Thanh Chau · Journal of Communications · 2020
Speech separation plays an important role in a speech-related system because it can denoise, extract and enhance speech signal, and after all improve the accuracy and performance of the system. In recent years, many approaches only separate the speech out of commonly high-frequency noise or a particular background sound. We propose a more powerful approach, combining an autoencoder and a bandpass filter to separate speech signals. This combination can extract the speech in the mixture with not only high-frequency noise but also many kinds of different background sounds. Our approach can be flexibly applied for the new background sounds. Experimental results show that our model can extract fastly and effectively the speech signal with 9.01 dB in SIR and 11.26 in SDR. On the other hand, we can adjust the passband to identify the range of frequency at the output signal to apply for particular applications.