Topic Modelling and Density-based Clustering for Opinion Anomaly Detection on Forum Discussion Transcript
Peter Yudhistira, Gregorius Satia Budhi, Alvin Nathaniel Tjondrowiguno · Research Square · 2024
Abstract A directed discussion is a platform where community members may discuss specific topics whose scope is defined by the forum's facilitator. Several issues arose in a setting where each speaker's comment regarding the discussion topic must be noted for further manual and thorough analysis. Due to the sheer volume of the speech data and the limitation of available manpower, the assigned annotator for a given discussion frequently lost focus and became unable to take accurate note of each speaker's comment, resorting to paraphrasing. The paraphrasing caused a loss of information, context, and personal bias that prevented an objective evaluation. This study intended to address the issue and is conducted to objectively retrieve the comments contributed by each discussion member regarding a given topic and utilise unsupervised learning to find and identify anomalous comments. In particular, comments that are not appropriate to the topic of the discussion become the focus of this study. To this end, we propose a feature extraction and anomaly detection pipeline. The obtained text corpus will have its features extracted with word embedding and topic modelling. The extracted features will then become input for the anomaly detection methods: density-based clustering (DBSCAN), Isolation Forest (IF), and Local Outlier Factor (LOF). This research aims to search for effective combinations of parameters to obtain said anomalies. Our results indicate that topic modelling with Latent Dirichlet Allocation (LDA) as a feature extraction method and DBSCAN could isolate inappropriate, unrelated comments with up to 71.42% accuracy.