Evaluation of keyness metrics: Reliability and interpretability

Lukas Sönning · 2022

While keyword analysis has become an essential tool in corpus-based work, the question of how to quantify keyness has been subject to considerable methodological debate. This has given rise to a variety of computerized metrics for detecting and ranking candidate items based on the comparison of a target to a reference corpus. Building on previous work, the present paper starts out by delineating four dimensions of keyness, which distinguish between frequency- and dispersion-related perspectives and identify substantively different aspects of typicalness. Existing measures are then organized according to these dimensions and evaluated with regard to two specific criteria, their interpretability and reliability. The first of these, which has been neglected in previous work, is a critical feature if metrics are to offer informative indications of keyness. The second criterion is performance-oriented and reflects the degree to which a metric produces stable and replicable rankings across repeated studies on the same pair of text varieties. Our illustrative analysis, which deals with the identification of key verbs in academic writing, shows considerable differences among indicators with regard to these two criteria. Our findings provide further support for the superiority of the Wilcoxon rank sum test and allow us to identify, within each dimension of keyness, metrics that may be given preference in applied work in light of our criteria.

Read the paper · More papers on PaperTik