Smooth soft mel-spectrographic masks based on blind sparse source separation
Marco Kühne, Roberto B. Togneri, Sven Erik Nordholm · 2007
Abstract This paper investigates the use of DUET, a recently proposedblind source separation method, as front-end for missing dataspeech recognition. Based on the attenuation and delay estima-tion in stereo signals soft time-frequency masks are designedto extract a target speaker from a mixture containing multiplespeech sources. A postprocessing step is introduced in order toremove isolated mask points that can cause insertion errors inthe speech decoder. The results for connected digit experimentsin a multi-speaker environment demonstrate that the proposedsoft masks closely match the performance of the oracle maskdesigned with a priori knowledge of the source spectra. Index Terms : speech recognition, missing data, attenuationand delay estimation 1. Introduction The concept of time-frequency (TF) masking has recently at-tracted some interest in the field of blind signal separation (BSS)[1, 2]. Demixing via TF-masks has the potential to separatemixtures with more sources than sensors as it does not rely onmatrix inversion. Instead the TF-plane is partitioned into dis-joint regions each assigned to a particular source. The sourcesare then recovered by converting each region back into the timedomain. It seems promising to use BSS systems as front-endsfor automatic speech recognition (ASR). In [3] we have pro-posed such a combination using a BSS technique called DUETand a missing data (MD) speech recognizer. The proposedsystem uses DUET to estimate TF-masks in the sparse Short-Time-Fourier-Transform (STFT) domain before converting thehigh STFT frequency resolution to a perceptual mel-frequencyscale suitable for ASR. In this way we can avoid source re-construction and directly exploit the spectrographic masks forMD-ASR. This paper extends our previous work in two regards.Firstly, we replace binary masks with soft masks which havebeen proven to be beneficial for both speech recognition andspeech enhancement. Several studies [2, 4] have reported thatbinary TF-masking can lead to audible unnatural sound arti-facts that can be avoided to some degree by soft masks. Theadvantages of soft masks in MD-ASR are even more evidentas marginalization approaches based on soft decisions consis-tently outperformed hard masks [5]. Secondly, we show thata simple mask postprocessing can lead to substantial recogni-tion improvements. A two-dimensional (2-D) median filter wasapplied to reduce the influence of outliers visible as scattered