Statistical measures of lexical associations

Douglas Biber, Susan Conrad, Randi Reppen · Cambridge University Press eBooks · 1998

The simplest way to identify collocate pairs is by their relative frequency – that is, by how commonly one pair, such as “large number,” occurs relative to another pair, such as “large man.” Such frequency information can give a sense of the most common collocational associations. However, frequency information alone may present a biased measure of the strength of associations between words. More common words are more likely to occur in a collocate pair simply by chance. Therefore, an alternative way to judge the strength of associations between words is to use measures that account for the likelihood of words occurring together by chance – i.e., statistical measures. Statistical measures can be used to analyze both the associations between collocate pairs and the differences between the collocations of particular words. Below we briefly review a common statistical test for each of these purposes. We explain the principles behind the tests, but for details about the statistical formulas, you should consult the articles listed under “Further reading.” Mutual information score The mutual information score or mutual information index gives a measure of the strength of association between two words. It focuses on the likelihood of two words appearing together within a particular span of words (the span is specified for the analysis, e.g., adjacent words, a window of three words, etc.).

Read the paper · More papers on PaperTik