Automatic Keyphrase Extraction based on NLP and Statistical Methods

Martin Dostál, Karel Ježek · 2011

In this article we would like to present our exper imental approach to automatic keyphrase extraction based on statistical methods and Wordnet-based pattern evaluation. Automatic keyphrases are import ant for automatic tagging and clustering because manually assigned keyphrases are not sufficient in most cases. Keyphrase candidates are extracted in a new way derived from a combination of graph methods (TextRank) and statistical methods (TF*IDF). Keyword candidates are merged with named entities and stop words according to NL POS (Part Of a Speech) patterns. Automatic keyphrases are generated as TF*IDF weighted unigrams. Keyphrases describe the main ideas of documents in a human-readable way. Evaluation of this approac h is presented in articles extracted from News web sites. Each article contain s manually assigned topics/categories which are used for keyword evalua tion.

Read the paper · More papers on PaperTik