Research on News Keyword Extraction Technology Based on TF-IDF and TextRank

Yao Lu, Pengzhou Zhang, Zhang Chi · 2019

With the rapid development of information technology and the widespread use of the Internet, the Internet, as an information carrier, has gradually replaced the traditional media such as newspapers and television, and become the main channel for people to obtain information. This paper takes English news text as the research object of keyword extraction method. We combine TF-IDF and the TextRank algorithm to extract keywords from text by constructing word graph model, counting word frequency and inverse document frequency, and considering the weight of the positioning of headlines. A large number of experiments have been carried out with Sina News Corpus, and the performance of the algorithm is evaluated by recall rate, precision rate and macro average value. The results show that the integration of TF-IDF and the TextRank algorithm significantly outperforms the traditional algorithm in performance parameters and extraction effect.

Read the paper · More papers on PaperTik