Leveraging the Multilingual Indonesian Ethnic Languages Dataset In Self-Supervised Models for Low-Resource ASR Task
Sakriani Sakti, Benita Angela Titalim · 2023
Indonesia is home to roughly 700 languages, which amounts to about ten percent of the global total, positioning it as the second-most linguistically diverse country after Papua New Guinea. Unfortunately, the availability of language technologies for languages in Indonesia is quite scant. Among those few technologies that have been developed, they mainly focus on the official Indonesian language, while a large number of Indonesian ethnic languages remain uncovered. To accelerate the development of language technology for under-resourced languages, especially for those languages in Indonesia, we contribute a multilingual Indonesian ethnic languages dataset for evaluation by the Multilingual Speech processing Universal PERformance Benchmark (ML-SUPERB). In this study, we also investigate low-resource automatic speech recognition (ASR) tasks in those languages on various self-supervised models, various minimum sets of training data, and various mismatched conditions.