Using Word Frequency Lists to Measure Corpus Homogeneity and Similarity between Corpora

Adam Kilgarriff · 1997

How similar are two corpora? A measure of corpus similarity would be very useful for language engineering. Word frequency lists are cheap and easy to generate so a measure based on them would be of use as a quick guide in many circumstances; for example, to judge how a newly available corpus related to existing resources, or how easy it might be to port an NLP system designed to work with one text type to work with another. Corpus similarity can only be interpreted in the light of corpus homogeneity. The paper presents a measure, based on the Ø 2 statistic, for measuring both corpus similarity and corpus homogeneity. A method for evaluating the accuracy of a corpus-similarity measure is introduced. The Ø 2 -based measure is compared with a rank-based measure and shown to outperform it. 1 Introduction How similar are two corpora? The question arises on many occasions. Does it matter whether language researchers use this corpora or that, or are they similar enough for it to make no ...

Read the paper · More papers on PaperTik