Newsgroup topic extraction using term-cluster weighting and Pillar K-Means clustering
Sigit Adinugroho, Randy Cahya Wihandika, Putra Pandu Adikara · International Journal of Computers and Applications · 2020
Topic extraction is an essential tool to help gathering information from a vast amount of sources. This paper introduces a new approach to extract topics from a collection of text documents. In order to obtain the topics, preprocessing steps are conducted to remove unnecessary parts of the documents. Then, a term frequency-inverse document frequency is built to weight terms in documents. After that, SVD-based feature transformation is involved in building features used for clustering. Prior to clustering process, the Pillar algorithm is run to select initial centroids for K-Means clustering. Finally, weights of terms in clusters are calculated using term-cluster weight as a basis to choose topics from clusters. Based on the experimental result, it is concluded that the framework achieves satisfactory results by attaining the accuracy of 100%, 95.1%, 83,7%, and 68.7% for 4 topics obtained from Binary2, Multi5, Multi7, and Multi10 categories of 20Newsgroup dataset.