Development of a multilingual isolated digits speech corpus

Emmanuel Malaay, Michael Simora, Ronald John Cabatic, Nathaniel Oco, Rachel Edita Roxas · 2017

We present a multilingual speech corpus for isolated digits. As case study, we focused on languages in the Philippines: English, Filipino, Ilocano, Cebuano, and Spanish. Our isolated digits speech corpus has a duration of almost nine hours, collection from 262 speakers. These data were word- level annotated and will be used to train the acoustic models using the ASR toolkits. The corpus will be used for an automatic speech recognition (ASR) system and therefore the database must be sufficient to develop an ASR system.

Read the paper · More papers on PaperTik