Nict-Tib1: A Public Speech Corpus Of Lhasa Dialect For Benchmarking Tibetan Language Speech Recognition Systems

Kak Soky, Zhuo Gong, Sheng Li · 2022

The Lhasa dialect is the primary Tibetan dialect, with the most speakers in Tibet and the most extensive written scripts over its lengthy history. Studying speech recognition methods in the Lhasa dialect significantly conserves Tibet’s distinctive linguistic variety. Previous research on Tibetan speech recognition focused on academic research on non-public datasets, e.g., selecting phone-level acoustic modeling units and incorporating tonal information, but had less contribution to limited data for the community. To solve the low-resource data problem, we introduce the NICT-Tib1 (phase1) database, a new open-sourced database for the Lhasa dialect. We further update benchmark systems under the monolingual and multilingual settings, respectively. Experimental results show that the performances of these models are consistent with previous work. We believe our work will promote the existing speech recognition research on the Tibetan language, and other low-resource languages.

Read the paper · More papers on PaperTik