Compounding in a Swedish Blog Corpus
Robert Östling, Mats Wirén · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2013
Research in compounding for Swedish has a long tradition at Stockholm University, with Benny Brodda starting already in 1967/68 (Brodda, 1981, p. 102). One of the sources that he used was an electronic version of SAOL 9, the Swedish Academy word list in its 9th edition (1950), with compound borders indicated. However, in later work he also looked solely at the forms of words and syllables without any lexical resource at hand (Brodda, 1981). Good compound analysis is highly needed for unrestricted text, especially for languages whose orthographies concatenate compound components (that is, juxtapose the components without an intervening space). This means that that every such concatenation corresponds to a word. This way of forming words is extremely productive in most Germanic languages (including Swedish, but with the exception of English) and, for example, Finnish, Hungarian and Greek, and provides an important reason why an exhaustive list of words remains impossible to construct in these languages. Also, in languages like this an unknown word will most likely be a compound (Stymne and Holmqvist, 2008). On a related note, compounding seems to be an area where a lot of the creativeness of language is put to work (Svanlund, 2009; De Smedt, 2012). So what new is there to say about Swedish compounding that could not be said one or a couple of decades ago? To begin with, there has been an enormous increase in the amount of electronically available data. At the Department of Linguistics, we have collected corpora from several Internet sources during the last years, including 2.7 billion tokens of Swedish blog text. There are two reasons why we find working with data from the Internet in general and blogs in particular highly useful. First, the sheer amount of data means that we obtain new ways of studying systematically various marginal and low-frequency phenomena that previously were more or less out of reach. One such example concerns neologisms and creative compounding in Swedish, which we can find by looking among words that are extremely low-frequency in spite of the large data set. Secondly, as has frequently been pointed out, text found on the web often has a colloquial and spontaneous character. The umbrella term here is user-generated content, that is, text published predominantly by non-professionals in media such as blogs, forums, reviews, social networks and wikis. User-generated content, including blogs, can thus provide an effective window into language change