Fast NMF based approach and improved VQ based approach for speech recognition from mixed sound

Shoichi Nakano, Kazumasa Yamamoto, Seiichi Nakagawa · Asia-Pacific Signal and Information Processing Association Annual Summit and Conference · 2012

We have considered a speech recognition method for mixed sound, consisting of speech and music, that removes only the music based on vector quantization (VQ) and non-negative matrix factorization (NMF). This paper describe fast calculation technique of music removal based on NMF and improvement using a VQ method. For isolated word recognition using the clean speech model, an improvement of 46% word error reduction rate was obtained compared with the case of not removing music. Furthermore, a high recognition rate, close to clean speech recognition was obtained at 10 dB. For the case of the multi-conditions, our proposed method reduced the error rate of 50% compared with the multi-conditions model.

Read the paper · More papers on PaperTik