Extracting Keyphrases from Chinese News Articles Using TextRank and Query Log Knowledge

Weiming Liang, Changning Huang, Mu Li, Bao‐Liang Lu · Institutional Repositories DataBase (IRDB) · 2009

Abstract. Keyphrases extracted from articles are beneficial in helping people boost brows-ing speed, but unfortunately keyphrases are rarely available for news articles due to the high expense of labor and time for manual annotation. This paper proposes a practical approach to extracting keyphrases for Chinese news articles using the TextRank and query log knowl-edge. Previous work is word based, while our approach uses phrase as its basic element. We generate phrases by employing several statistical criteria with the huge amount of queries as a training corpus. We use TextRank, a graph-based learning algorithm, for extracting keyphrases from Chinese news articles. In addition, two instructive features, lengths and positions of phrases, are incorporated into the TextRank model. Experimental results demon-strate that our methods improve the performance significantly.

Read the paper · More papers on PaperTik