Advancing towards automatic speech recognition of children innaturalistic preschool environment
Satwik Dutta, Dwight Irvin, John H. L. Hansen · The Journal of the Acoustical Society of America · 2021
The preschool classroom is a viable space for capturing young children’s interactions with teachers. Developing speech/language assessment tools will offer a non-invasive way to quantify child-teacher interactions, allowing teachers to make changes as needed and better support children’s speech, language, and cognitive skills. Building Automatic Speech Recognition (ASR) systems for child speech has been challenging, since most studies have been conducted: (1) with older children, (2) in clean and controlled settings, and (3) generally are limited to prompted or read stimuli. To date, limited research has focused on developing ASR systems for spontaneous child speech in preschool settings. For our study, spontaneous child-teacher conversations were captured in preschool classrooms using the LENA digital audio recorder. A 25-h corpus of child speech were used for experiments using an open-source ASR toolkit. Major challenges in this data for developing ASR models include noisy background environments, age and diversity of children. Our effort focuses on alternative time-delay neural networks (TDNN) for acoustic modelling to improve ASR, as well as trade-offs in language model development and lexicon/vocabulary partitioning. Results from this study across acoustic and language models will be presented using child preschool classroom data. [Work sponsored by NSF CyberLearning Grant #1918032.]