Reconnaissance de phones fondée sur du Transfer Learning pour des enfants apprenants lecteurs en environnement de classe

Lucile Gelin, Morgane Daniel, Thomas Pellegrini, Julien Pinquier · HAL (Le Centre pour la Communication Scientifique Directe) · 2020

Current performance of speech recognition for children is below that of the state-of-the-art for adultspeech. Young children’s speech is particularly difficult to recognise, and substantial corpora aremissing to train acoustic models. Furthermore, in the scope of our reading assistant for 5-7-year-oldchildren learning to read, models need to cope with slow reading rate, disfluencies, and classroom-typical babble noise. In this paper, we compare acoustic models for phone recognition on child speechusing data that is very noisy and limited in quantity. We show that transfer learning with adult-trainedtime-delay neural networks and three hours of child speech improves the phone error rate by 7.6%relative, over a model trained on child speech. The addition of vocal tract length normalisation onadult speech further reduces the error rate by 5.1% relative, reaching a PER of 37.1%.

Read the paper · More papers on PaperTik