Keyword Extraction from Scientific Articles in Bahasa Indonesia using TextRank Algorithm

Dani Gunawan, Fanindia Purnamasari, Ranti Ramadhiana, Romi Fadillah Rahmat · 2020

Some of text mining application usually uses the keyword extraction method to obtain the text features. Most of the text mining application implements keyword extraction in the pre-processing stage. The keyword extraction method is a necessary process to obtain the depiction of an article or document. In general, this research is divided into three major stages. First of all, the pre-processing stage. This stage prepares and cleans the scientific article from insignificant parts (such as header, footer, page number, author name, or author affiliation) before the keyword extraction stage. Second, the keyword extraction stage applies the TextRank algorithm to extract the keywords. Third, the post-processing stage to generate keyphrase based on the rank of the vertices and their adjacent. The results show that the number of assigned keywords plays a crucial role in determining the recall value. The recall value of the TextRank algorithm raises from 38.46% (5 keywords) to 55.26% (10 keywords) and 61.54% (15 keywords). It occurs to the keyword extraction from scientific articles with reference inclusion. A similar result occurs to the keyword extraction from scientific articles without references inclusion. Future research can consider eliminating the post-processing stage by including multiword expression candidates in the pre-processing stage to improve keyword extraction time.

Read the paper · More papers on PaperTik