Mongolian speech corpus IMUT-MC

Liu Zhiqiang Liu Zhiqiang, MA Zhiqiang MA Zhiqiang, ZHANG Xiaoxu ZHANG Xiaoxu, BAO Caijilahu BAO Caijilahu, XIE Xiulan XIE Xiulan, ZHU Fangyuan ZHU Fangyuan · Science Data Bank Datasets · 2022

Mongolian, as a minority language, lacks a large-scale speech corpus accessible to researchers for experimental support, because its users are scattered and it is difficult to collect and label the speech sounds, which hinders the further development of Mongolian speech recognition. This research group has constructed a speech corpus IMUT-MC for Mongolian speech recognition tasks, which contains about 212 hours of reading speech recorded by 417 speakers, and is committed to advancing Mongolian speech recognition research.

Read the paper · More papers on PaperTik