Guiding Ideas

Łukasz Dębowski · 2020

This chapter focuses on the expertise of mathematicians. It explores some aspects of ideal statistical language models as seen via quantitative laws exhibited by texts spontaneously created by humans. The chapter investigate some particular statistical patterns of natural texts that may inform us of more general theoretical properties of human language. Speaking of concrete statistical patterns of language, it begins with Zipf’s law. Zipf’s law, discovered by Estoup and later popularized by Zipf and Mandelbrot, is undoubtedly the most famous of all quantitative patterns exhibited by natural texts known so far – for an overview of many other patterns see Kohler et al. Herdan’s law describes a relationship between the number of word types V, i.e. the number of different words in some initial part of a text, and the number of word tokens N, i.e. the number of all words in the same part.

Read the paper · More papers on PaperTik