A Graphical Model for Formant Tracking
James Malkin, Xiao Li, Jeffrey A. Bilmes · 2006
We present a novel approach to estimating the first two formants (F1 and F2) of a speech signal using graphical models. Using a graph that takes advantage of less commonly used features of Bayesian networks, both v-structures and soft evidence, the model presented here shows that it can learn to perform reasonably without large amounts of training data, even with minimal processing on the initial signal. It far outperforms a factorial HMM using the same assumptions and suggests that with further refinement the model may produce high quality formant tracks.