Developing Filipino and Cebuano ASR Systems with Children’s Speech Corpora on Healthcare Monitoring

Ronald M. Pascual, Gretel Mendoza, Alessandro Miguel Zapanta, Kyle Angelo Lino · 2025

Automatic Speech Recognition (ASR) systems play a crucial role in converting spoken language from audio recordings into text, facilitating the extraction of linguistic data, and the creation of transcriptions. These can be applied in healthcare where audio serves as an alternative input to applications like a chatbot. Creating ASR systems for these purposes presents obstacles, particularly for low-resource languages such as Filipino and Cebuano in the Philippines. Although a previous study has effectively trained ASR models for Filipino and Cebuano using healthcare-oriented datasets with adult speakers, the difficulty persists in using these systems to decode children’s speech. This study aims to bridge this gap by developing baseline Filipino and Cebuano ASR systems trained on children’s speech and tailored for healthcare monitoring purposes. Additional children’s speech data were collected and processed while existing base corpora from a healthcare chatbot project were reviewed. Various versions of the dataset and configurations were also used to experiment with lowering the word error rate (WER). The best-performing iteration of the Filipino ASR model had a word error rate of 25.59% while the Cebuano ASR performed significantly better with a 16.80% word error rate. These performed reasonably well when compared to Google ASR. Future works involve increasing the data, using more thorough and complex preprocessing techniques, and adjusting the ASR system to better cater to healthcare specific use.

Read the paper · More papers on PaperTik