Latvian Speech Corpus (LaRKo)
Ilze Auziņa, Roberts Darģis, Kristīne Levāne-Petrova, Kristīne Pokratniece, Daira Vēvere · 2014
The Latvian Speech Corpus (around eight hours of audio data) consists of the broadcasts that have appeared in the mass media, including the audio recordings of the meetings of the Latvian Saeima and their orthographic transcripts. Each audio recording is accompanied by meta data about the place and duration of the recording, as well as the gender and approximate age of the speaker.