In search of optimal data selection for training of automatic speech recognition systems
A. Nagroski, Lou Boves, Herman J. M. Steeneken · 2004
This paper presents an extended study in the topic of optimal selection of speech data from a database for efficient training of ASR systems. We reconsider a method of optimal selection introduced in our previous work and introduce variosearch as an alternative selection method developed in order to find a representative sample of speech data with a simultaneous control of acoustical and statistical parameters of data selected. Next, we present experiments in which the performance of a standard ASR system trained with data sets selected from a Dutch digits database via different selection methods was compared. The results show that the length of training utterances has a dominant impact on the recognition performance. Therefore, the length of the utterances is a factor that must be taken into account when interpreting phoneme recognition scores.