Cross Lingual Sentiment Analysis using Modified BRAE
Sarthak Jain, Shashank Batra · 2015
Cross-Lingual Learning provides a mech-anism to adapt NLP tools available for la-bel rich languages to achieve similar tasks for label-scarce languages. An efficient cross-lingual tool significantly reduces the cost and effort required to manually an-notate data. In this paper, we use the Recursive Autoencoder architecture to de-velop a Cross Lingual Sentiment Analysis (CLSA) tool using sentence aligned cor-pora between a pair of resource rich (En-glish) and resource poor (Hindi) language. The system is based on the assumption that semantic similarity between different phrases also implies sentiment similarity in majority of sentences. The resulting sys-tem is then analyzed on a newly developed Movie Reviews Dataset in Hindi with la-bels given on a rating scale and compare performance of our system against exist-ing systems. It is shown that our approach significantly outperforms state of the art systems for Sentiment Analysis, especially when labeled data is scarce. 1