Recent Latvian Speech Corpora for Linguistic Research and Technology Development
Ilze Auziņa, Normunds Grūzītis, Roberts Darģis, Guna Rābante-Buša, Didzis Goško, Jānis Vempers, Raivis Kivkucāns, Artūrs Znotiņš · Baltic Journal of Modern Computing · 2024
This paper presents newly created Latvian speech corpora aimed at advancing both linguistic research and speech technology development.Although multilingual models like XLS-R and Whisper have reduced the amount of data needed for fine-tuning speech recognition models even for less-resourced languages, diverse and curated speech corpora remain essential.We provide an overview of several recent Latvian speech corpora, emphasizing their importance for both general-purpose and domain-specific use cases and comparing their design with previously created speech datasets for Latvian.We also introduce a common platform for analysing open-access Latvian speech corpora, and discuss initial evaluation and integration of speech recognition models fine-tuned on the new datasets for practical speech transcription and post-editing applications in research and industry.Finally, we present a competitive open-source speech recognition model for Latvian.