Implementation of Robust Speech Recognition by Simulating Infants' Speech Perception Based on the Invariant Sound Shape Embedded in Utterances

Nobuaki Minematsu, Susumu Asakawa, Yu Qiao, Daisuke Saito, T. Nishimura, Kenzo Hirose · 2009

Recently, a novel and structural representation of speech was proposed [1, 2], where inevitable acoustic biases caused by static extra-lingusitic factors are completely removed from speech. This speech structure is composed of only transform-invariant (topologically invariant) speech contrasts or dynamics [3] with no use of absolute and static acoustic features such as spec-trums. Although this framework posed a problem of so strong invariance that two different words could be evaluated as the same, our previous study successfully introduced good constraints to the invariant features [4]. We realized the invariance only with respect to speaker variability through Multiple Stream Struc-turalization (MSS) [4]. In this paper, after introduction of our proposed representation, we describe that it can be a goodmath-ematicalmodel of infants ’ speech perception [5, 6, 7, 8], percep-tual constancy of speech [9, 10, 11, 12], and Jakobson’s classi-cal theory of relational invariance [13, 14, 15]. Next, we show new experimental results using the proposed speech representa-tion. The high robustness is verified again by using frequency warped utterances, which simulate the utterances of very tall speakers and very short ones. Further, we also investigate this representation using a phoneme-balanced word set because, in the previous study, we used only an artificial word set comprised of vowel sequences such as /aeoui/. Results are very promising. 1.

Read the paper · More papers on PaperTik