A dataset of Tibetan-Chinese speech translation
Xiaobing ZHAO, Jialuo LIU, Mengting Zhou, Xue Feng Jiang, Xiaoke Qi · China Scientific Data · 2024
The advancement of research frontiers in speech translation relies upon the quality and diversity of available datasets. Currently, the exploration of speech translation for minority languages is subject to numerous constraints due to the limited availability of publicly accessible dataset datasets. To address this gap, this paper aims to construct and release a dataset of speech translation from Tibetan speech to Chinese text. The dataset is derived from the WeChat public platform and publicly available Tibetan speech recognition datasets. We collected the data with the assistance of web scraping and machine translation and performed manual segmentation and annotation. And the data underwent expert review and correction to ensure its accuracy and quality, resulting in a high-quality dataset of Tibetan-Chinese speech translation. The dataset comprises 7,270 entries with a total size of 965 MB. The dataset can not only provide a foundational data framework for exploring Tibetan-to-Chinese speech translation, but also contribute to the advancement of relevant technologies and algorithms. Moreover, it is expected to offer substantial support for the application of speech translation systems in the context of minority languages.