A Mispronunciation-Based Voice-Omics Representation Framework for Screening Specific Language Impairments in Children
Wei Bo, Matthew Rubino, Wenyao Xu · 2024
This paper introduces an innovative end-to-end (E2E) framework for screening Specific Language Impairment (SLI) in children, centralizing phoneme-level mispronunciation (PLM) detection to enhance the precision and reliability. We have developed a unique voice-omics representation that translates PLM predictions into symbolic sequences, yielding significant phenotyping biomarkers that provide objective and quantifiable assessments of children's speech patterns. Through meticulous fine-tuning of the Connectionist Temporal Classification (CTC) model on the L2-ARCTIC dataset and rigorous five-fold cross-validation, our E2E models have demonstrated remarkable ac-curacy, with Area Under the Curve (AUC) values exceeding 0.71 and a notable recall rate of up to 71.5 % on the CHILDES dataset. Our approach signifies a substantial advancement in SLI screening, leveraging cutting-edge technology to capture the complexities of spontaneous speech in children.