Empirical and Theoretical Bases of Zipf's Law
Ronald E. Wyllys · Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign) · 1981
1Let us start by considering a basic form of Zipf's law. Suppose one has a natural-language corpus, e.g., a book written in English. Next, suppose one makes a frequency count of the words in the corpus, i.e., counts the number of occurrences of the, and, of, etc. Finally, suppose one arranges the words in decreasing order of frequency so that the most frequent word has rank 1; the next most frequent, rank 2; and so on. For example, a frequency count of the 75 word-types (i.e., dictionary entries) represented by the 112 word-tokens (i.e., distinct occurrences) in the two preceding paragraphs yields the partial results shown in table 1. This set of rank-ordered frequency counts, though quite small for the purpose, serves moderately well as an illustration of the fact that rank and frequency have a surprisingly constrained relationship in natural-language corpora. The values of the products of rank r and frequency f fall in the relatively limited range 27-30 in the middle of table 1, and we may note that there was no a priori reason for us to expect that the middle products rf would fall within so limited a range.