The effects of speaker training on ASR accuracy
Stephen Anderson, Natalie Liberman, Larry Gillick, S. Foster, Sahoko Hama · 1999
In our experiment, 30 computer-literate elderly speakers (15 male, 15 female) with no previous ASR experience were given 2 hours of intensive training in using a speech recognition system. Before and after this training session, they were asked to read separate 520-word texts. Measuring the word error rates (WERs) on these “before training” and “after training” recordings, we find a small but statistically significant improvement. Before training, speakers had an average WER of 20.9%, and after training, 19.8%. We examine changes in speaking rate, phrase length, and SNR and their impact on WER.