Mixture models for word frequency distributions
Fiona J. Tweedie, R. Harald Baayen · 2000
Word frequency distributions are generally extremely skewed and are described as having a Large Number of Rare Events (LNRE). LNRE distributions have been found to provide fits to many examples of word frequency distributions. However, Baayen and Tweedie (1998b) present a distribution of the Dutch suffix -heid which cannot be fitted by standard methods. In this case, the data come from a composite source and we introduce the idea of mixture distributions to deal with this. We present expressions for the expected word frequency distribution, the expected number of tokens, and the number of types in the population. An acceptable fit to the -heid data is presented.