Split size‐rank models for the distribution of index terms
Michael J. K. Nelson, Jean M. Tague · Journal of the American Society for Information Science · 1985
Abstract Since the introduction of the Zipf distribution, many functions have been suggested for the frequency of words in text. Some of these models have also been applied to the distribution of index terms in a set of documents. The models are of two forms: rank‐frequency and frequency‐size. The former serve well to describe the distribution of high‐frequency terms; the latter the distribution of low‐frequency terms. In this article, a split model is proposed, which uses both a rank function for the high frequency terms and a size function for the low frequency terms, with the point of transition being determined either empirically or by rule. This model is fitted to the marginal empirical term distributions for four document datasets. Distributions to describe index term exhaustivity and term co‐occurrence are also considered briefly.