An integrated approach to the detection and classification of accents/dialects for a spoken document retrieval system
Sharmistha Gray, John H. L. Hansen · 2005
In this study, an integrated approach to accent/dialect detection and classification is proposed, which can be used for enhancing Rich indexing of historical spoken documents with accent/dialect information. A next generation spoken document retrieval (SDR) system would require a more diverse set of speech criteria including speaker, accent/dialect, language, stress/emotion and environment content. The proposed accent/dialect tagging system for SDR is based on several recent advances in a multi-dimensional space. Here, temporal and spectral based features including the stochastic trajectory model (STM), pitch structure, formant location and voice onset time (VOT) are considered. Mono-phone based STM (MP-STM) is shown to be the most successful for dialect classification with an average rate of 96.5% for read speech and 72.5% for spontaneous speech, while classifying four dialects. An example of next generation Rich transcript indexing for conversational speech to simulate SDR is also presented.