Non-linear frequency warping for VTLN using subglottal resonances and the third formant frequency

Harish Arsikere, Steven M. Lulich, Abeer A. Alwan · 2013

This paper proposes a non-linear frequency warping scheme for VTLN. It is based on mapping the subglottal resonances (SGRs) and the third formant frequency (F3) of a given utterance to those of a reference speaker. SGRs are used because they relate to formants in specific ways while remaining phonetically invariant, and F3 is used because it is somewhat correlated to vocal-tract length. Given an utterance, the warping parameters (SGRs and F3) are determined by obtaining initial estimates from the signal, and refining the estimates with respect to a speaker-independent model. For children (TIDIGITS), the proposed method yields statistically-significant word error rate (WER) reductions (up to 15%) relative to conventional VTLN (linear warping) when: (1) speakers show poor baseline performance, and/or (2) training data are limited. For adults (Wall Street Journal), the WER reduction relative to conventional VTLN is 4-5%. Comparison with other non-linear warping techniques is also reported.

Read the paper · More papers on PaperTik