Japanese vowel recognition using external structure of speech
Takao Murakami, Kazutaka Maruyama, Nobuaki Minematsu, Kenzo Hirose · 2005
Speech acoustics is inevitably distorted by non-linguistic features such as vocal tract length, gender, age, microphone, room, line, hearing characteristics, and so on. Recently, a novel acoustic representation of speech was proposed, called the acoustic universal structure [1, 2]. It discards all the absolute properties of speech events and captures only their interrelations or contrasts to represent external structure of speech. Based on a mathematical model of the non-linguistic distortions, the new representation can remove the inevitable distortions as cepstrum smoothing can remove pitch information from speech. Recognition experiments using the external structure were conducted and it was shown that the proposed structure models trained from a single speaker with no normalization can outperform the conventional speaker-independent HMMs with CMN and SS. It was also found that lowpass filtering and white noise addition as preprocessing improved the performance because they can suppress the distortions rather well