Noise robust speech activity detection

Waleed Habib Abdulla, Zhou Guan, Hou Chi Sou · 2009

An efficient noise robust feature is presented to track the speech activity in noisy environments. Speech is modeled by one class of 16 phone-like Gaussian mixtures while noises are modeled by 15 classes of 6 mixtures each. The feature vector used is a concatenation of carefully selected coefficients from MFCC, LPCC, and their first and second derivatives. A finite state machine and energy validation components are proposed as post-processor for the GMM classifier to rectify the misclassified speech segments. The demonstrated speech activity detection system based on our feature detects reliably both speech and non-speech segments. The designed frame work has been benchmarked against the commercially available codecs G.729, GSM-EFR, MR1, and MR2. Results show the proposed technique outperforms all these commonly used techniques under various SNR levels and in different noisy environments.

Read the paper · More papers on PaperTik