Sentiment analysis using a random forest classifier on turkish web comments
PERVAN Nergis KELEŞ · Communications Faculty of Sciences University of Ankara Series A2-A3 Physical Sciences and Engineering · 2017
Sentiment analysis is an activeresearch area since early 2000s as a field of text classification. Most of thestudies in this field focus on the analysis using the text in English language,where the Turkish and the other languages have fallen behind. The purpose ofthis research is to contribute to the text analysis in Turkish language usingthe contents that we access through web sites. In particular, we deduce thesentiment behind noisy product reviews and comments in a highly popularcommercial web page. In this context, we generate a unique dataset that includes9100 product review samples for training our classification model. There aredifferent word representation methods that are utilized in sentiment analysis,such as bag-of-words and n-gram models. In this work, we generated our wordmodels using the word2vec algorithm. In this model, each word in the vocabularyis represented as a vector of 300 dimensions. We utilize 70% of our dataset inthe training of a Random Forest Model and make binary classification ofsentiments as being positive or negative, utilizing the ratings of the user forthe product as classification labels. In the highly noisy and unfilteredcomments, we achieve an accuracy of 84.23%.