Multi-lingual label alignment using acoustic-phonetic features derived by neural-network technique
Paul Dalsgaard, Ove Andersen, William J. Barry · 1991
In previous work on label alignment, encouraging results were obtained using selected acoustic-phonetic features to model the individuals speech phonemes. Selection was based on minimal covariance between features on the one hand, and the inclusion of features underlying critical phonological opposition on the other. In the present work, principal component analysis was applied to give a number of uncorrelated output parameters which maximally exploit the discriminatory power of the features and are derived independently of the phonological functionality. Results of label alignment on three different European languages, Danish, English, and Italian, using different numbers of principal parameters show that the accuracy with ten parameters is at least as good as with 15 manually selected features. The best result is found for British English, which has 78% of its phoneme transition boundaries positioned within +or-20 ms of manually placed reference boundaries.>