Adaptation of acoustic model for Indonesian using varying ratios of spontaneous speech data
Devin Hoesen, Dessi Puji Lestari, Masayu Leylia Khodra · 2016
This paper presents our work in determining the ratio/amount of speech adaptation data that gives the optimum recognition error rate. Ten triphone acoustic models are first built in a similar-to-10-fold cross-validation manner using the dictated speech corpus. The whole dictated speech is read from 10 prepared transcripts by a diverse group of 301 Indonesian speakers. Each round of training uses utterances from one of the prepared transcripts. The resulting triphone models are also evaluated against its corresponding spontaneous speech evaluation set. The models that yield the lowest, the highest, and the closest-to-mean recognition error are then adapted to its corresponding spontaneous speech adaptation data using the Maximum A-posteriori Probability (MAP) method. The amount of the corresponding spontaneous speech adaptation data for each model is varied from 10% to 100% with a 10% increment. Thus for each of the 3 triphone models, there will be 10 resulting adapted models. The resulting adapted models are evaluated against their corresponding spontaneous and dictated speech evaluation set. The trend in the results shows that around 30-60% spontaneous speech adaptation data (roughly translates to 11.5 to 24 hours of speech) gives the most optimum (low) recognition error rate.