Noise-robust automatic speech recognition with exemplar-based sparse representations using multiple length adaptive dictionaries
Emre Yılmaz, Jort Florent Gemmeke, Hugo Van hamme · Lirias · 2013
In this work, we apply our recently proposed sparse representations based speech recognition system on the small vocabulary track of the 2 nd ‘CHiME ’ Speech Separation and Recognition Challenge. This system uses exemplars of different length to approximate noisy speech segments as a linear combination of the speech and noise exemplars with sparse weights. The exemplars are labeled speech segments extracted from the training data, each representing half words and they are organized in multiple dictionaries based on their class and length. A reconstruction error-based decoding is adopted to find the best matching class sequence. After the initial experiments on AURORA-2, we further apply our system on the CHIME data which is a more challenging task addressing not only non-stationary noise but also reverberation. Moreover, the structure of the CHIME data allows speaker-dependent acoustic modeling and sampling noise segments from the immediate acoustic context of the target utterances. Using speaker-dependent dictionary sets, several recognition experiments are conducted on the development and test sets to evaluate the system performance with different kinds of noise dictionaries. These experiments show that combined noise dictionaries containing noise exemplars extracted from both the immediate acoustic context of the test utterances and noise-only segments in the training data provide better recognition accuracies compared to fixed and adaptive dictionaries. Index Terms — Exemplar-based recognition, sparse representations, non-negative sparse coding, multiple dictionaries 1.