Biphone-rich versus Tripohne-rich: A Comparison of Speech Corpora in Automatic Speech Recognition
Yong-Chang Yio, Min-Siong Liang, Yuang-Chin Chiang, Ren-Yuan Lyu · 2005
In this paper, we compare the performance of a speech recognition system trained with two speech corpora. We select two set of words such that they covered all the cross-syllable bi-phones and tri-phones, and are called phonetically biphone-rich and triphone-rich respectively. It is required about 10 times more words than that of cross-syllable biphones to cover all the cross-syllable triphones. To facilitate fair comparison, the biphone-rich corpus is thus consisted often sets of words that each covers all the cross-syllable biphones. With those words as data sheets, a male Taiwanese speaker recorded all the words as microphone speech. The resulting speech corpora, about 100 minutes for each set, are used to train for the acoustic models. Although both perform quite well in tasks with recognition networks of linear net and free syllable net, the triphone-rich corpus does not show much advantages over the biphone-rich corpus.