An improvement to TF-IDF: Term Distribution based Term Weight Algorithm

Tian Xia, Yanmei Chai · Journal of Software · 2011

In the process of document formalization, term weight algorithm plays an important role. It greatly interferes the precision and recall results of the natural language processing(NLP) systems. Currently, TF-IDF term weight algorithm is widely applied into language models to build NLP Systems. Since term frequency is not the only discriminator which is necessary to be considered in term weight ing and make each weight suitable to indicate the term’s importance, we are motivated to investigate other statistical characteristics of terms and found an important discriminator: term distribution. Furthermore, we found that , in a single document, a term with higher frequency and close to hypo-dispersion distribution usually contains much semantic information and should be given higher weight. One the other hand, in a document collection, the term with higher frequency and hypo-dispersion distribution usually contains less information. Based on this hypothesis, by leveraging the Pearson Chi-square Test Statistic, a Term Distribution based Local Term Weight Algorithm and Global Term Weight Algorithm are put forward respectively in this paper. Also, the experiment results at the end of this paper approve the reliability and efficiency of the algorithm s .

Read the paper · More papers on PaperTik