Amdo-Chinese Speech Translation Dataset

Zengzhuoma Ren, Liping ZHU, Xiaobing ZHAO, Ning Li · Data Intelligence · 2024

The advancement of speech translation research, particularly for minority low-resource languages, is impeded by the scarcity of publicly available datasets. This paper addresses this challenge by introducing and releasing the Amdo Tibetan-Chinese Speech Translation Dataset (AMDO-CHS AST). The audio recordings in this dataset encompass the Amdo dialect spoken in the Aba pastoral area in Sichuan Province, the Gannan pastoral area in Gansu Province, and the Qinghai pastoral area, capturing speech data from individuals of diverse genders aged between 12 and 35 years. Text data undergoes machine translation to generate initial translations, subsequently refined through meticulous proofreading by skilled professionals. Following preprocessing steps such as resampling and normalization, we compile a dataset comprising 11 hours and 6091 pairs of data, with an average audio duration of 6.51 seconds. The creation of this dataset establishes a foundational resource for advancing research in Amdo Tibetan-Chinese speech translation.

Read the paper · More papers on PaperTik