Speech recognition based on inferred articulatory movements
George J. Papcun, Timothy R. Thomas · The Journal of the Acoustical Society of America · 1990
A neural network was trained to learn the associations between speech acoustics, suitably transformed and normalized, and associated articulatory movements, which were recorded at the University of Wisconsin x-ray microbeam facility. Training was done on the numbers “one” through “ten” spoken by a female speaker; articulatory templates were selected from her speech. Speech from a male speaker was transformed and normalized and passed through the trained neural network. Articulatory movements inferred from his speech by the neural network were compared, by Euclidian distance, with the templates from the speech of the female speaker. To assign a single point of match for each template, an extremum finding algorithm was passed over the function that relates temporal position to degree of match. A threshold, defined as a distance of more than two standard deviations from the mean of the distance measure taken at each step in the speech sample, was used to successfully recognize the words spoken. Enhanced discrimination was obtained by weighting the articulatory parameters. [Work supported by DOE Contract W-740S-ENG-36.]