Automatically Building a Corpus for Sentiment Analysis on Indonesian Tweets

Alfan Farizki Wicaksono, Clara Vania, Bayu Distiawan, Mirna Adriani · Institutional Repositories DataBase (IRDB) · 2014

The popularity of the user generated content, such as Twitter, has made it a rich source for the sentiment analysis and opinion mining tasks. This paper presents our study in automatically building a training corpus for the sentiment analysis on Indonesian tweets. We start with a set of seed sentiment corpus and subsequently expand them using a classifier model whose parameters are estimated using the Expectation and Maximization (EM) framework. We apply our automatically built corpus to perform two tasks, namely opinion tweet extraction and tweet polarity classification using various machine learning approaches. Experiment result shows that a classifier model trained on our data, which is automatically constructed using our proposed method, outperforms the baseline system in terms of opinion tweet extraction and tweet polarity classification.

Read the paper · More papers on PaperTik