Voice Activity Detection. Fundamentals and Speech Recognition System Robustness
Javier Ramı́rez, Jijomon Chettuthara Moncy, José David Cely Callejas · 2007
Robust Speech Recognition and Understanding 2 recently reported strategies by assessing the speech/non-speech discrimination accuracy and the robustness of speech recognition systems. ApplicationsVADs are employed in many areas of speech processing.Recently, VAD methods have been described in the literature for several applications including mobile communication services (Freeman et al. 1989), real-time speech transmission on the Internet (Sangwan et al., 2002) or noise reduction for digital hearing aid devices (Itoh and Mizushima, 1997).As an example, a VAD achieves silence compression in modern mobile telecommunication systems reducing the average bit rate by using the discontinuous transmission (DTX) mode.Many practical applications, such as the Global System for Mobile Communications (GSM) telephony, use silence detection and comfort noise injection for higher coding efficiency.This section shows a brief description of the most important VAD applications in speech processing: coding, enhancement and recognition. Speech codingVAD is widely used within the field of speech communication for achieving high speech coding efficiency and low-bit rate transmission.The concepts of silence detection and comfort noise generation lead to dual-mode speech coding techniques.The different modes of operation of a speech codec are: i) the active speech codec, and ii) the silence suppression and comfort noise generation modes.The International Telecommunication Union (ITU) adopted a toll-quality speech coding algorithm known as G.729 to work in combination with a VAD module in DTX mode. Figure 1 shows a block diagram of a dual mode speech codec.The full rate speech coder is operational during active voice speech, but a different coding scheme is employed for the inactive voice signal, using fewer bits and resulting in a higher overall average compression ratio.As an example, the recommendation G.729 Annex B (ITU, 1996) uses a feature vector consisting of the linear prediction (LP) spectrum, the fullband energy, the low-band (0 to 1 KHz) energy and the zero-crossing rate (ZCR).The standard was developed with the collaboration of researchers from France Telecom, the University of Sherbrooke, NTT and AT&T Bell Labs and the effectiveness of the VAD was evaluated in terms of subjective speech quality and bit rate savings (Benyassine et al., 1997).Objective performance tests were also conducted by hand-labeling a large speech database and assessing the correct identification of voiced, unvoiced, silence and transition periods.Another standard for DTX is the ETSI (Adaptive Multi-Rate) AMR speech coder (ETSI, 1999) developed by the Special Mobile Group (SMG) for the GSM system.The standard specifies two options for the VAD to be used within the digital cellular telecommunications system.In option 1, the signal is passed through a filterbank and the level of signal in each band is calculated.A measure of the SNR is used to make the VAD decision together with the output of a pitch detector, a tone detector and the correlated complex signal analysis module.An enhanced version of the original VAD is the AMR option 2 VAD, which uses parameters of the speech encoder, and is more robust against environmental noise than AMR1 and G.729.The dual mode speech transmission achieves a significant bit rate reduction in digital speech coding since about 60% of the time the transmitted signal contains just silence in a phone-based communication.