Suprasegmental features and continuous speech recognition
Pierre Dumouchel · 2002
We first propose to model microprosody by means of a Bayesian classifier assuming multivariate Gaussian distributions on suprasegmental features. Second, we normalize the suprasegmental features by using dynamic parameters extracted from diphones. Third, we examine three different types of covariance matrices and show that a full covariance matrix per diphone gives the best results. Finally, the insertion of the microprosodic model into the INRS large vocabulary speech recognition improves the word recognition slightly from 48% to 52%.>