Comparison of Sentiment Analysis of Review Comments by Unsupervised Clustering of Features Using LSA and LDA

Shu-Cih Tseng, Yu-Ching Lu, Goutam Chakraborty, Long‐Sheng Chen · 2019

Text documents could be classified using words as features. As the number of words in the vocabulary is large, the dimension of the document space will be very high. In that case, the feature vector for a document is too long, and very sparse, and it makes clustering and classification algorithms fail. There are various ways to reduce this dimension. In this work, we used Latent Semantic Analysis (LSA), which is actuated by Singular Value Decomposition (SVD). After SVD, we have a compact representation of the documents, which are clustered. In a separate experiment, we did topic modeling using Latent Dirichlet Allocation (LDA). In this initial work, our premise is that comments are of two categories, positive and negative. We cluster the document, in the reduced dimensional space, into two, using K-means clustering. After dimension reduction by LSA and LDA, the ground truth for documents in two clusters was verified manually, and the results compared.In this work, we used tourists' comments as documents. Tourists visit to a place is influenced by comments from previous visitors. Our final goal is to extract factors that lead to positive comments and those leading to negative comments. That would help promoting tourist business by focusing on the factors that really matters for the customers.

Read the paper · More papers on PaperTik