Latvian Speech Corpus (LaRKo)

Ilze Auziņa, Roberts Darģis, Kristīne Levāne-Petrova, Kristīne Pokratniece, Daira Vēvere · 2014

The Latvian Speech Corpus (around eight hours of audio data) consists of the broadcasts that have appeared in the mass media, including the audio recordings of the meetings of the Latvian Saeima and their orthographic transcripts. Each audio recording is accompanied by meta data about the place and duration of the recording, as well as the gender and approximate age of the speaker.

Read the paper · More papers on PaperTik