Pronunciation Adaptation For Disordered Speech Recognition Using State-Specific Vectors of Phone-Cluster Adaptive Training
R. Sriranjani, Srinivasan Umesh, M. Ramasubba Reddy · 2015
Pronunciation variation is a major problem in disordered speech recognition.This paper focus on handling the pronunciation variations in dysarthric speech by forming speaker-specific lexicons.A novel approach is proposed for identifying mispronunciations made by each dysarthric speaker, using state-specific vector (SSV) of phone-cluster adaptive training (Phone-CAT) acoustic model.SSV is low-dimensional vector estimated for each tied-state where each element in a vector denotes the weight of a particular monophone.The SSV indicates the pronounced phone using its dominant weight.This property of SSV is exploited in adapting the pronunciation of a particular dysarthric speaker using speaker-specific lexicons.Experimental validation on Nemours database showed an average relative improvement of 9% across all the speakers compared to the system built with canonical lexicon.