From data mining to sentiment analysis : Classifying documents through existing opinion mining methods
Ville Jukarainen · Theseus (Ammattikorkeakoulujen) · 2012
This thesis proposes a solution for document-level opinion mining, a method of finding overall opinion from given sources, for example, product reviews, news articles and blogs. This suggestion was done by using existing methods and an unsupervised self-organizing map for classification. The task is to create a system that can classify documents written in the English language, according to opinion categories, for example, positive, neutral and negative. Also, a design suggestion is made for how the presented solution could be implemented in Cluetail Ltd. systems, using Python programming-language. Thesis process started by learning about the underlying techniques (Machine learning, data mining and natural language processing). These techniques create the foundation for learning sentiment analysis. Partially supervised and unsupervised learning methods were chosen for this approach. Lexicon based positive - negative term appearance features, and bigram features extracted according to part-of-speech tags and with calculated opinion orientation using an unsupervised statistical meth-od. These features formed a feature vector for each document which describes the found overall opin-ion. Two review datasets with known opinion categories were used, and the capabilities were tested both with opinion polarities (positive - negative) and multiple opinion categories (very positive, posi-tive, neutral, negative, very negative). The findings from these results, issues, improvements and the ways to extend this work in other languages are discussed on the last pages of this thesis. Overall this work is the first step towards functional document-level opinion mining system, and a simple introduction to sentiment analysis meant for anyone who may have interest to learn about it.