Comparing the nonstandard language of Slovene, Croatian and Serbian tweets
Darja Fišer, Tomaž Erjavec, Nikola Ljubešić, Maja Petrović · Repository of the University of Ljubljana (University of Ljubljana) · 2015
In this paper we carry out a cross-lingual comparison of nonstandard features in the language of social media for Slovene, Croatian and Serbian. The goal of the analysis is twofold: (1) we try to establish the extent to which the observed phenomena are universal rather than language-specific, and (2) we propose an approach for automatic scoring of (non)standardness levels of user-generated content, which can be used as a separate annotation layer in corpora. Quantitative and qualitative analyses of the results show that the majority of the language used on Twitter is fairly standard, especially in Slovene and Croatian. The prevalent characteristic of nonstandard Slovene tweets is nonstandard orthography, while nonstandard lexis is more typical of Serbian tweets, possibly due to a younger user profile.