Combination of Latent Dirichlet Allocation (LDA) and Term Frequency-Inverse Cluster Frequency (TFxICF) in Indonesian text clustering with labeling

Lya Hulliyyatus Suadaa, Ayu Purwarianti · 2016

Due to the limited labeled data, clustering is a solution for classifying documents that do not have prior knowledge. The combination of Latent Dirichlet Allocation (LDA) in grouping documents by topic and Term Frequency-Inverse Cluster Frequency (TFxICF) in the labeling was proposed to resolve the problem of classification using clustering completed with a description of the cluster results. Indonesian text preprocessing has been done by the extraction of abbreviations and acronyms, tokenization, stemming and stopwords elimination. Experiments were conducted using Indonesian digital library documents, 113 documents from digital library of STIS and 60 documents from digital library of ITB, to examine the effects of text preprocessing, to compare the cluster results of LDA with other clustering algorithms and to compare the use of word and phrase tokens in the clustering and labeling. The cluster quality was measured by using precision, recall, and F-measure and the label quality was determined by similarity with the keywords that most frequently appear in the clusters. Based on the experimental results, preprocessing techniques can improve the cluster quality. LDA algorithm produces documents cluster by topic with cluster quality better than K-Means and Lingo. Word based LDA generates cluster with better quality than phrase based LDA. Moreover, the labeling by using word based TFxICF is more descriptive than phrase based TFxICF. Therefore, the use of word based LDA for clustering and phrase based TFxICF for labeling was proposed.

Read the paper · More papers on PaperTik