Creating and Working with Corpus of Spoken Lithuanian
Kamandulyt edot -Merfeldien edot Laura, Godliauskas Povilas · Frontiers in artificial intelligence and applications · 2014
In this article we give a general overview on the development of the Corpus of Spoken Lithuanian. We consider the methodology of the corpus as well as the process of transcribing and coding the collected data. In addition, we provide a brief analysis based on the collected corpus data of spoken adult speech (ADS). The analysis is conducted according to four aspects: distribution of parts of speech, inflectional changes, features of syntax, and usage of loanwords. In the end of the paper we provide conclusions and future expectations.