Voice activity detection in noisy environments

Jan Stadermann, Volker Stahl, Georg Rose · 2001

ABSTRACTThe subject of this paper is robust voice activity detec-tion (VAD) in noisy environments, especially in car en-vironments. We present a comparison between severalframe based VAD feature extraction algorithms in com-bination with different classifiers. Experiments are car-ried out under equal test conditions using clean speech,clean speech with added car noise and speech recordedin car environments. The lowest error rate is achievedapplying features based on a likelihood ratio test whichassumes normal distribution of speech and noise and aperceptron classifier. We propose modifications of thisalgorithm which reduce the frame error rate by approxi-mately 30% relative in our experiments compared to theoriginal algorithm.1. INTRODUCTIONA voice activity detector (VAD) is an algorithm whichis able to distinguish between speech (usually distortedby noise) and noise only. The output from a VAD is asignal that possesses the information whether the inputsignal contains speech (e.g. output value 1 ) or noiseonly (e.g. output value 0). In order to make the problemmore tractable we assume that the speech and the noisesignal are stationary within a certain time interval. Thisassumption allows to apply conventional techniques ofsignal processing to this problem.The most common features used by VAD algorithms arerelated to signal energy. Since the frame energy aloneshows bad performance,a signal-to-noise ratio (SNR) isintroduced that uses an estimation of the noise energy.Further improvement is achieved if the SNR computa-tion is done separately for every spectral componentandif probabilitydensitiesare introducedforthe spectralen-ergy values. Other algorithms use features like the zerocross rate or the autocorrelationfunctionto find discrim-inative features in the time domain. The classification isdone by comparing results from the feature extractionwith an adaptive threshold [1] or using classification al-gorithms from pattern recognition[2]. The algorithms in[3] and [4] combineseveral features to detect speech. [5]post-processes the VAD decision with so-called “hang-over” methods. These methods use context (i.e. the clas-sification result of previous frames) in order to achievea more reliable results. This paper reviews some VADfeature extraction algorithms described in the literatureand combines them with classification algorithms com-monly used in pattern recognition.We proposea modifi-cation of one of the feature extraction algorithms, whichimproved the frame error rate by 30% relative. The pa-per is organized as follows: Section 2 describes our ba-sic VAD setup and the most important VAD algorithmstogether with modifications, Section 3 summarizes theused classifiers, Section 4 presents results, and conclu-sions are given in Section 5.2. VAD ALGORITHMSAll VAD algorithms considered in this paper are framebased. The incoming audio signal is sampled, quantizedand divided into overlapping frames, then each frameis classified as either speech or non-speech. Throughoutour experiments we used a frame length of

Read the paper · More papers on PaperTik