Auditory-based acoustic-phonetic signal processing for robust continuous speech recognition

Ahmed Ali, Jan Van der Spiegel · Scholarly Commons (University of Pennsylvania) · 1999

State-of-the-art automatic speech recognition (ASR) systems are significantly inferior to humans especially in the presence of noise and other adverse conditions. Substantial research into both the human auditory system and the acoustic-phonetic characteristics of speech and their variability is needed to build improved front-end processing systems capable of extracting the useful, information-rich, acoustic features. This is expected to have a profound effect on the automatic speech recognition systems whose performance and robustness can improve by integrating more acoustic-phonetic knowledge into their design. This work examines several auditory-based speech processing systems for their formant extraction ability from clean and noisy speech. A novel speech processing system, based on the biological Average Localized Synchrony Detection (ALSD) principle, is presented. The new system compares favorably with other auditory-based systems in formant extraction. The acoustic-phonetic characteristics of speech are explored with the goal of extracting features that capture the phonemic identities of sounds independent of the sources of variability of context, gender, dialect, and noise. Minimal sets of static, dynamic, spectral and temporal features are proposed for various phoneme recognition tasks. Novel knowledge-based, hard- and soft-decision, algorithms are developed to extract, combine and manipulate multiple features in the decision-making process. The recognition accuracy obtained exceeds that of any known knowledge-based work, which shows a clear enhancement in the acoustic-phonetic knowledge. The obtained accuracy is also comparable with that of statistical classifiers in similar tasks. Moreover, the extracted features are shown to be robust in the presence of noise. This work makes original contributions to research in auditory-based speech processing, the acoustic-phonetic characteristics of speech, knowledge-based ASR, and robust ASR. It is also expected to be a step in the direction of building a hybrid ASR system that combines substantial acoustic-phonetic knowledge in statistical frameworks like Hidden Markov Models (HMMs).

Read the paper · More papers on PaperTik