The Hot Keyphrase Extraction Based on TF*PDF

Yan Gao, Jin Liu, PeiXun Ma · 2011

Key phrase consisting of several words is viewed as the phrase that represent the topic and the content of the whole text. Extracting key phrase is a good way to detect hot topics and tracking topics from news report. In this paper, a two-step key phrase extraction method based on TF*PDF is proposed. In the first step, the position-weighted TF*PDF algorithm is proposed to obtain candidate hot term list and the bursty value of term is used to filter the noise in the list. In the second step, a phrase identification process combines hot terms into phrases using position information, frequency information etc. At last the position-weighted TF*PDF algorithm are also used to weight the phrase, and the top k phrases are chosen as hot key phrases. The experiments on the real web data indicate that our extraction method provides solutions with improved quality.

Read the paper · More papers on PaperTik