The 2nd CHiME Speech Separation and Recognition Challenge: Approaches on Single-Channel Speech Separation and Model-Driven Speech Enhancement

Pejman Mowlaee, Juan Andrés Morales Cordovilla, Franz Pernkopf, Hannes Pessentheiner, Martín Hagmüller, Gernot Kubin · 2013

In this paper, we address the small vocabulary track (track 1) described in the CHiME 2 challenge dedicated to recognize utterances of a target speaker with small head movements. The utterances are recorded in a reverberant room acoustics corrupted with highly non-stationary noise sources. Such adverse noise scenario imposes a challenge to state-of-the-art automatic speech recognition systems. We developed two individual front end sf or the output of the delay-and-sum beamformer: (i) a model-driven single-channel speech enhancement stage which combines the knowledge of the speaker identity modeled by a trained vector quantizer with a minimum statistics based noise tracker, an d( ii) as ingle-channel source separation stage which employs models of the target speaker as well as the background noise as codebooks. Our perceived signal quality and separation results averaged on the CHiME 2 development set justify the effectiveness of both strategies in terms of recovering the target speech signal. Also, our best results on keyword recognition accuracy show 20% improvement over the provided baseline results on the development and test sets.

Read the paper · More papers on PaperTik