A Preliminary Study on Applying the Conditional Modeling to Automatic Dialect Classification
Rongqing Huang, John H. L. Hansen · 2006
This paper addresses the advances in unsupervised dialect classification. There are no transcripts for both the training data and the testing data. In this study, we view the classification problem in speech in an recognition-based way instead of the conventional generative model-based approach and try to bypass the unknown transcript problem. The new algorithm is based on conditional model. The new algorithm has two notable advantages: first, it can train a statistical model without transcripts, so it can work in our transcript-free classification problem; second, the conditional model in the new algorithm can allow arbitrary feature representations, therefore, it can encode more discriminative features than the generative models such as hidden Markov model (HMM), which has to use the independent and local features due to the model restrictions. The conditional model used in the study is the conditional random fields (CRF). Further study on combining the generative model and conditional model is presented. In the Spanish dialect classification evaluation, the CRF and the combined modeling technique show some interesting results