Extended Minimum Classification Error Training in Voice Activity Detection

Takayuki Arakawa, Haitham Hassanieh, Masanori Tsujikawa, Ryosuke Isotani · 2009

Voice activity detection (VAD) is a fundamental part of speech processing. Combination of multiple acoustic features is an effective approach to make VAD more robust against various noise conditions. There have been proposed several feature combination methods, in which weights for feature values are optimized based on minimum classification error (MCE) training. We improve these MCE-based methods by introducing a novel discriminative function for whole frames. The proposed method optimizes combination weights taking into account the ratio between false acceptance and false rejection rates as well as the effect of the use of shaping procedures such as hangover.

Read the paper · More papers on PaperTik