Dialect/Accent Classification via Boosted Word Modeling
Rongqing Huang, John H. L. Hansen · 2006
The paper addresses novel advances in English dialect/accent classification/identification. A word level based modeling technique is proposed that is shown to outperform a LVCSR based system with significantly less computational cost. The new algorithm, which is named WDC (word-based dialect classification), converts the text independent decision problem into a text dependent problem and produces multiple combination decisions at the word level rather than make a single decision at the utterance level. There are two sets of classifiers employed for WDC, word classifier, D/sub W(k)/, and utterance classifier, D/sub u/. D/sub W(k)/ is boosted via the real AdaBoost.MH algorithm in the probability space directly instead of the feature space. D/sub u/ is boosted via the dialect dependency information of the words. Two dialect corpora are used in the evaluation. Significant improvement in dialect classification is achieved for both corpora.