Utilization of the Redundancy of the Phonemic Transcription of Speech for Automatic-Speech Recognition

Klaus W. Otten, Richard T. Kleiner · The Journal of the Acoustical Society of America · 1964

An automatic-speech-recognition system segments continuous speech into phonemes and identifies those phonemes. This segmentation and acoustical-recognition process is far from perfect. A high-order linguistic-recognition process, which considers the restrictions in the use of phonemes in sequences, can correct a large percentage of the errors made in the acoustical-recognition process. A study and simulation program was conducted to determine the relationship between the redundancy of the spoken langauge in its phonemic transcription and the potential percentage-error correction as a function of the original error rate. Models of 4, 8, and 16 phonemes with redundancies comparable to those of spoken languages were synthesized. These synthesized phoneme sequences were artificially distorted. Through use of first-order conditional probabilities, partial correction of the distorted sequences was achieved. The results of the computer-simulation program indicate that speech recognition with low error rates is feasible even if the initial acoustical recognition has a high error rate, provided that the statistical properties of the acoustical-recognition elements are known and that they are used in a linguistic error-correcting recognition process as simulated. The simulation program and the numerical results are presented.

Read the paper · More papers on PaperTik