Decision combination in speech metadata extraction

Xiaofan Lin · 2004

Speech metadata extraction can both improve speech recognition and enable novel interactive voice response applications. Unlike the previous research, which concentrates on the frame-level signal processing and pattern classification, this paper systematically studies the behavior of decision combination at the utterance level. We analyze the asymptotic characteristics, and the factors affecting frame-level classification. In addition, we introduce new methods to more accurately and efficiently combine frame-level decisions, including phoneme/power-based weighting and smart sampling. Experimental results in gender classification are presented.

Read the paper · More papers on PaperTik