Character Expressions in Meta-Learning for Extremely Low Resource Language Speech Recognition
Rui Zhou, Akinori Ito, Takashi Nose · 2024
For the construction of a high-quality speech recognition system, a substantial volume of annotated speech data is requisite. However, preparing such expansive datasets is impracticable for a vast majority of global languages. Therefore, we need to develop speech recognition systems for low-resource languages. In this paper, we propose a method where the model is initially pretrained on speech data from ten languages utilizing meta-learning. After this, the model undergoes fine-tuning using a small amount of speech data of the target language. We used only about 15 minutes of speech and acheved CER of less 30%. Although the set of alphabet differs from language to language, alphabets often show phonetic resemblances. Capitalizing on this observation, we proposed the method of alphabet unification, employed English pronunciation rules based on Latin alphabet as a standard to align the alphabetic representation across languages under study. Our results show that such alphabet unification approach enhanced the performance.