Scaling laws and fluctuations in the statistics of word frequencies
Martin Gerlach, Eduardo G Altmann · New Journal of Physics · 2014
In this paper, we combine statistical analysis of written texts and simple stochastic models to explain the appearance of scaling laws in the statistics of word frequencies. The average vocabulary of an ensemble of fixed-length texts is known to scale sublinearly with the total number of words (Heaps' law). Analyzing the fluctuations around this average in three large databases (Google-ngram, English Wikipedia, and a collection of scientific articles), we find that the standard deviation scales linearly with the average (Taylorʼs law), in contrast to the prediction of decaying fluctuations obtained using simple sampling arguments. We explain both scaling laws (Heaps' and Taylor) by modeling the usage of words using a Poisson process with a fat-tailed distribution of word frequencies (Zipfʼs law) and topic-dependent frequencies of individual words (as in topic models). Considering topical variations lead to quenched averages, turn the vocabulary size a non-self-averaging quantity, and explain the empirical observations. For the numerous practical applications relying on estimations of vocabulary size, our results show that uncertainties remain large even for long texts. We show how to account for these uncertainties in measurements of lexical richness of texts with different lengths. Introduction and background . A characteristic signature of complex systems is the appearance of scaling laws, e.g. heavy-tailed distributions or allometric scaling, offering a unifying perspective on seemingly unrelated phenomena irrespective of the microscopic details of the underlying process. Multiple scaling laws frequently appear in the same system, in which case a major interest is to find the quantitative relationship between them. For instance, a heavy-tailed distribution in the frequency of items is directly related to the sub-linear scaling of the number of different items with sample size. Main results . We report a new scaling law in the statistics of words in written texts (fluctuation scaling or Taylorʼs law) and show how to relate it to other scaling-laws known to exist in these systems (see figure 1 ). By modeling the usage of words by a simple stochastic process we show that all scaling laws appear simultaneously only if topical variations across different texts are considered. Wider implications . Owing to the generality of our analytical approach, our results can be applied to other complex systems in which similar scalings hold, e.g. ecology, allometric scaling of cities, or network growth. Furthermore, our analysis allows for a quantification of the uncertainties around these scaling laws and suggests more appropriate null models to assess the validity of scaling laws in empirical data. Figure. Three different scaling laws observed in empirical data of word frequencies (English Wikipedia). (a) Zipfʼs law: frequency of the -th most frequent word. (b) Heaps' law: the number of different words, , as a function of text-length, , for each individual article (black dots). (c) Taylorʼs law: standard deviation, , as a function of the mean, , for the vocabulary conditioned on the text-length (computed over different articles). Poisson (dark line) shows the expectation from a Poisson null model assuming the empirical rank-frequency distribution from (a). (Data: ) (pale line) shows the mean, , and standard deviation, , of the data within a running window in . For comparison, we show in (c) the scalings and (dashed).