Histogram Equalization for Robust Speech Recognition
Luz Garca, José Carlos, N. de la Torre, Carmen Bentez, Antonio J. · InTech eBooks · 2008
Speech Recognition, Technologies and Applications 24•Quick adaptation for non-native speakers.Current voice applications demand robustness and adaptation to non-native speakers' accents.• Databases with realistic degradations.Formulation, recording and spreading of voice databases containing realistic examples of the degradation existing in practical environments are needed to face the existing challenges in voice recognition.This chapter will analyze the effects of additive noise in the speech signal, and the existing strategies to fight those effects, in order to focus on a group of techniques called statistical matching techniques.Histogram Equalization -HEQ-will be introduced and analyzed as main representative of this family of Robustness Algorithms.Finally, an improved version of the Histogram Equalization named Parametric Histogram Equalization -PEQ-will be exposed. Voice feature normalization Effects of additive noiseWithin the framework of Automatic Speech Recognition, the phenomenon of noise can be defined as the non desired sound which distorts the information transmitted in the acoustic signal difficulting its correct perception.There are two main sources of distortion for the voice signal: additive noise and channel distortion.Channel distortion is defined as the noise convolutionally mixed with speech in the time domain.It appears as a consequence of the signal reverberations during its transmission, the frequency response of the microphone used, or peculiarities of the transmission channel such as an electrical filter within the A/D filters for example.The effects of channel distortion have been fought with certain success as they become linear once the signal is analyzed in the frequency domain.Techniques such as RASTA filtering, echo cancellation or Cepstral mean subtraction have proved to eliminate its effects.Additive noise is summed to the speech signal in the time domain and its effects in the frequency domain are not easily removed as it has the peculiarity to transform speech nonlinearly in certain domains of analysis.Nowadays, additive noise constitutes the driving force of research in ASR: additive white noises, door slams, spontaneous overlapped voices, background music, etc.The most used model to analyze the effects of noise in the oral communication (Huang, 2001) represents noise as a combination of additive and convolutional noise following the expression