Automatic Keyphrase Extraction from Bengali Documents: A Preliminary Study

Kamal Krishna Sarkar · 2011

Key phrases are sequence of words that capture the main topics covered in a document. The key phrases help readers rapidly understand, organize, access and share information of a document. In this paper, we present a preliminary study on key phrase extraction from Bengali documents using two important features, such as TF*IDF, phrase's first occurrence in the text. For this study, we design a prototype system which works as follows: extracts n-grams from a source article, identifies candidate key phrases, and finally ranks the candidate key phrases to select the desired number of key phrases. The system has been tested on a collection of Bengali documents selected from a Bengali corpus downloadable from TDIL website and the preliminary results on Bengali key phrase extraction have been reported in this paper.

Read the paper · More papers on PaperTik