INESC-ID: Sentiment Analysis without Hand-Coded Features or Linguistic Resources using Embedding Subspaces
Ramón Fernández Astudillo, Silvio Amir, Ling Wang, Bruno Martins, Mário J. Silva, Isabel M. Trancoso · 2015
We present the INESC-ID system for the message polarity classification task of SemEval 2015.The proposed system does not make use of any hand-coded features or linguistic resources.It relies on projecting pre-trained structured skip-gram word embeddings into a small subspace.The word embeddings can be obtained from large amounts of Twitter data in unsupervised form.The sentiment analysis supervised training is thus reduced to finding the optimal projection which can be carried out efficiently despite the little data available.We analyze in detail the proposed approach and show that a competitive system can be attained with only a few configuration parameters.