Measures of Word Commonness
Petr Savický, Jaroslava Hlaváčová · Journal of Quantitative Linguistics · 2002
The main goal of this paper is to investigate methods of how to rank words in a way that corresponds to an intuitive notion of ‘commonness’. Since there is no formal definition of such a notion, our techniques may be considered as a suggestion for such a definition. The commonness of words is sometimes roughly substituted with their frequency in a language corpus. In order to suggest a better measure, we define a quantity, which we call corrected frequency. It depends not only on the frequency of a word in a corpus, but also on its distribution within the corpus. Unlike previous solutions of the same problem, we take the corpus as an uninterrupted sequence of words with no regard to borders between files, texts, genres, or any others. We introduce three different corrected frequencies. Their definitions are based on notions of information theory and analysis of random processes. Their values for individual words depend on the corpus. Hence, it is important to what extent they are stable with respect to the selection of the corpus. In order to investigate the suggested corrected frequencies from that point of view, we compare their values on five different subcorpora of the whole corpus. We present several examples of words taken from the Czech National Corpus that demonstrate in which way the corrected frequencies correspond to the intuitive commonness of these words.