A dataset of Mongolian-Chinese speech translation
qi xiao ke qi xiao ke, Borjigin B.Teniger Borjigin B.Teniger, Yuan Sun Yuan Sun, Xiaobing Zhao Xiaobing Zhao · Science Data Bank Datasets · 2022
Due to the lack of public datasets, few researches focus on speech translation in minority languages. To this end, this paper constructs a dataset of Mongolian-Chinese speech translation, named as NMLR-Mon2Chs ST. The dataset consists of Mongolian speech, Mongolian and Chinese text. First, Mongolian speech were obtained from 36 Mongols aged between 20 and 25 by recording on their mobile phones. Then, the corresponding Chinese texts were annotated by professionals. In order to make sure the quality of the dataset, the preprocessing was done, such as removing the quiet speech, resampling, and normalization. As a result, a total of 25 hours of high-quality data are obtained, and the average duration of audio in the dataset is 4.2 seconds. The establishment of this dataset allows researchers access to speech translation for minority languages.