Use of articulatory signals in automatic speech recognition
Louis D. Braida, Michael Picheny, Jordan R. Cohen, W. M. Rabinowitz, Joseph S. Perkell · The Journal of the Acoustical Society of America · 1986
Automatic speech recognition systems generally attempt to determine the spoken message from analysis of the acoustic speech waveform. In this research we evaluated the performance of the IBM Speech Recognition System [F. Jelinek, Proc. IEEE 73, 1616–1624 (1986)] when the input included measurements of selected articulatory actions occurring during speech production. The system achieved significant recognition rates (for isolated words in sentences) when the acoustical signal was disabled and the input was restricted to articulatory signal similar to those sensed by users of the Tadoma method of tactile-speech recognition [e.g., Reed et al., J. Acoust. Soc. Am. 77, 247–257 (1985)]. In other tests the availability of articulatory inputs improved recognition performance when the acoustical signal was sufficiently degraded by additive white noise. Improvements were observed independent of whether the recognition system made use of the likelihood that words would appear in the vicinity of the other words in a given sentence.