Automatic Speech Recognition using Pitch Information in Dynamic Bayesian Networks

Todd Andrew Stephenson, Mathew Magimai.-Doss, Hervé A. Bourlard · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 2000

The challenge of automatic speech recognition (ASR) increases when speaker variability is encountered. Being able to automatically use dierent acoustic models according to speaker type might help to increase the robustness of ASR. We present a system that attempts to do so by augmenting the standard acoustic observations with pitch information. This allows the system to use acoustic models more appropriate to speech with the given pitch. Furthermore, pitch information is more easily detected in noisy conditions; thus, it may be of use in robust speech recognition. Using dynamic Bayesian networks (DBNs) allows further renement of the system by eliminating unnecessary statistical dependencies and thus reducing the number of parameters. We show that when a system is trained on observed pitch data and performs recognition with missing pitch data, it can perform signicantly better than a system that uses acoustics information only.

Read the paper · More papers on PaperTik