M2ASR-MONGO: A Free Mongolian Speech Database and Accompanied Baselines
Tiankai Zhi, Ying Shi, Wenqiang Du, Guanyu Li, Dong Wang · 2021
Deep learning has significantly improved the performance of automatic speech recognition (ASR), in particular for major languages such as English and Chinese. However, for minor languages such as Mongolian, ASR performance is still limited, mostly due to the restricted amount of speech data. In this paper, we publish a free Mongolian speech database and the associated resources. The entire database involves 170 hours of speech signals produced by 259 native speakers. To our knowledge, this is the largest Mongolian speech database that is publicly available and free so far. We also publish two baseline recipes using the new database. The first one is based on Kaldi and aims to demonstrate the hybrid architecture (DNN-HMM), while the second one is based on Espnet and aims to demonstrate the end-to-end architecture (Transformer). This publication is a part of the M2ASR project, and all the resources are free for the research purpose.