Collocation and colligation

Tomas Lehecka · John Benjamins Publishing Company eBooks · 2015

of Firth's work).Since the terms were introduced, collocation in particular has become a fundamental concept in usage-based studies in many linguistic fields, most notably lexical syntax and semantics.Typically, collocations and colligations are studied in large electronic corpora which allows for statistical analyses of the co-occurrence patterns of linguistic items. CollocationCollocation refers to the syntagmatic attraction between two (or more) lexical items: morphemes, words, phrases or utterances.Most often, however, collocation analyses have been conducted on the word-level (see discussion in Hoey 2005: 158-159).The concept of collocation is based on the notion that each word in a language prefers certain lexical contexts over others, i.e. that any given word tends to co-occur with certain words more often than it does with others.For example, the word grass is often used together with green, and the lexeme LET T ER is often used together with the lexemes W RIT E A N D REA D (see e.g.Kjellmer 1996: 83).The strength of this kind of attraction between words can be measured through the statistical analysis of corpus data.The purpose of these statistical calculations is to find word pairs with significantly more co-occurrences than what would be expected by chance, given the words' total frequencies in the data.Thus, we can establish the most significant collocates of any given word in the language variety that the data represents (Sinclair 1966: 418, Berry-Roghe 1973: 103, Hoey 1991: 6-7).The syntagmatic attraction, or collocation strength, between two words W1 and W2 (a node and its collocate) is calculated based on four observed absolute frequencies in the data: (i) the total number of word tokens in the corpus, (ii) the number of tokens of W1 in the corpus, (iii) the number of tokens of W2 in the corpus, and (iv) the number of tokens where W1 and W2 co-occur within a specified distance from each other (see collocation window below).The observed number of co-occurrences in the corpus is compared to the expected number of co-occurrences, i.e. the number expected by chance given (i), (ii) and (iii).If the observed number of co-occurrences of W1 and W2 is larger than what can be ascribed to chance, then W2 is a statistically significant collocate of W1.

Read the paper · More papers on PaperTik