Towards reliable speech recognition in operating room noise environment
Petr Zelinka, Milan Sigmund · 2010
This paper describes several practical steps for accurate statistical modeling of a known acoustical noise environment to attain good performance of a small vocabulary speech recognizer for isolated words based on whole-word hidden Markov models. Hierarchical segmentation based on Bayes information criterion and k-means clustering followed by split-merge Gaussian mixture model training were utilized for noise model estimation. Parallel model combination technique produces final noise-corrupted speech models for a small group of speakers. Experiments were carried out on a real operating room ambient noise recorded during a neurosurgery at the University Hospital in Marburg.