A Bayesian approach to speaker normalization using vowel formant frequency
Dhananjay Ram, Debasis Kundu, Rajesh Mahanand Hegde · 2014
Large variation in speakers causes significant performance degradation of a speaker independent speech recognition system. In an attempt to compensate for this degradation in performance, this paper proposes a novel Bayesian approach to estimate speaker normalization parameters. An affine model is used here, which captures the variation in length of the vocal tract more effectively than the linear model used in literature. The vocal tract length normalization (VTLN) parameters are estimated using Least Squares Estimation (LSE) as well as a Bayesian approach which utilizes the Gibbs sampler, a special type of Markov Chain Monte Carlo method. Finally, a Mahalanobis distance based vowel recognizer is proposed and experiments are performed for both gender dependent and independent cases. Results clearly indicate a performance improvement for the Bayesian case over LSE.