A dataset of Mongolian-Chinese speech translation for the news field—M2CST-Mongo

Xiaobing ZHAO, Xue Feng Jiang, Jialuo LIU, Borjigin B.Teniger Borjigin B.Teniger, Xiaoke Qi · China Scientific Data · 2024

Datasets are the basis for training and evaluating speech translation systems, playing a vital role in driving innovative research in speech translation and facilitating progress in the field. However, the Mongolian-Chinese speech translation corpus is now relatively scarce, making it difficult with sufficient scale and diversity to support the training of translation models. As a result, the development of Mongolian-Chinese speech translation technology is faced with great challenges. In order to address this problem, we have constructed a dataset of Mongolian-Chinese speech translation for the news field in this study. First, we referred to previous research ideas on speech translation datasets to convert the publicly available Mongolian speech recognition dataset into a speech translation dataset. After data processing, the dataset is submitted for expert review and inspection. Through correction and analysis of this dataset, we produced a high-quality dataset of Mongolian-Chinese speech translation. This dataset is oriented to the news field, covering politics, economy, culture and other topics, with a total duration of 106.5 hours. There are 47,935 audio entries from 258 speakers with text in both Mongolian and Chinese. The total size of the dataset is 19.6 MB. The dataset fully takes into account the balance of letters to ensure the usability of the data. This dataset is expected to provide a certain data basis for exploring low-resource Mongolian-Chinese speech translation, promote the development of Mongolian-Chinese speech translation technology, and facilitate Mongolian-Chinese cultural exchanges.

Read the paper · More papers on PaperTik