Automatic text classification method based on Zipf’s law

V. A. Yatsko · Automatic Documentation and Mathematical Linguistics · 2015

This paper describes a method for automatic text classification based on analysing the deviation of the word distribution from Zipf’s law, combined with the zonal data-processing approach. Deviation is understood as the difference between the actual numerical score of a word and its score according to Zipf’s law. The proposed method involves the division of input and reference texts into J 0, J 1, and J 2 zones, and the creation of a numerical series using the words that are contained in the J 0 zone. The constructed numerical series shows the difference between the real scores of words and the scores calculated according to Zipf’s law. The proposed method can significantly reduce text dimensionality and thus improve the running speed of automatic text classification.

Read the paper · More papers on PaperTik