Sentiment Analysis on Monolingual, Multilingual and Code-Switching Twitter Corpora

David Vilares, Miguel Á. Alonso, Carlos Gómez‐Rodríguez · 2015

We address the problem of performing polarity classification on Twitter over different languages, focusing on English and Spanish, comparing three techniques: (1) a monolingual model which knows the language in which the opinion is written, (2) a monolingual model that acts based on the decision provided by a language identification tool and (3) a multilingual model trained on a multilingual dataset that does not need any language recognition step.Results show that multilingual models are even able to outperform the monolingual models on some monolingual sets.We introduce the first code-switching corpus with sentiment labels, showing the robustness of a multilingual approach.

Read the paper · More papers on PaperTik