Not just frequency

Stefan Τh. Gries · Amsterdam studies in the theory and history of linguistic science. Series 4, Current issues in linguistic theory · 2025

Abstract For decades, nearly all approaches to keyness analysis in corpus linguistics have been based on computing for each word type in question a single statistic — usually, the log-likelihood score G 2 — and ranking word types by how key that statistic made a word type for a target corpus T . In this paper, I discuss a new approach to keyness that (i) uses three dimensions of information (frequency in T , association to T , and dispersion in T relative to R and that (ii) measures both association and dispersion using the information-theoretic measure of the Kullback-Leibler divergence. I outline the computational steps and provide R code in a markdown document as well as a ready-made R function Keyness3D with which readers can conduct analyses of their own data. I exemplify the use of the function and its results using the learned text category in the Brown corpus against the rest.

Read the paper · More papers on PaperTik