Performance Comparison of Topic Modeling Algorithms on Indonesian Short Texts
Nuraisa Novia Hidayati, Anne Parlina · 2022
The number of short texts produced daily has increased significantly as a form of social communication commonly used on the internet. Extracting topics from extensive collections of short texts is one of the most challenging tasks in natural language processing, but it has numerous applications in the real world. The purpose of this study is to compare the topic extraction performance of the Latent Dirichlet Allocation (LDA), Non-Negative Matrix Factorization (NMF), and Gibbs Sampling Dirichlet Multinomial Mixture (GSDMM) algorithms from Indonesian short texts. The data was gathered from news articles about electric vehicles published on the online news site (Kompas.com). Regarding topic coherence scores, our results show that LDA outperforms NMF and GSDMM. However, human judgment indicates that the word clusters produced by NMF and GSDMM are easier to conclude.