Formant estimation of whispered speech based on spectral segmentation
Gong Chenghui, Zhao Heming, Gang Lü, Liu Jianxin · 2006
Whispered speech, without vocal cord vibration and always in low SNR, is more difficult both in its analysis and recognition. Thus its formant estimation becomes prominent in each field. The proposed algorithm is based on spectral segmentation. The complete spectrum is segmented into K segments, each of which contains a single formant. Here, improved dynamic programming and selective LP (linear predictive) methods are used. The former offers segment boundaries, and the latter leads to the parameters of formant frequency and its bandwidth as well. For whispered speech, the gain of vocal tract transfer function is also important. The tests are carried on Chinese whispered vowels, and the proposed algorithm is proved to be efficient. In low SNR, the segment based LP method is obviously superior to the conventional LPC and LSP