Unsupervised Word Usage Similarity in Social Media Texts

Spandana Gella, Paul Cook, Bo Han · Joint Conference on Lexical and Computational Semantics · 2013

We propose an unsupervised method for automatically calculating word usage similarity in social media data based on topic modelling, which we contrast with a baseline distributional method and Weighted Textual Matrix Factorization. We evaluate these methods against a novel dataset made up of human ratings over 550 Twitter message pairs annotated for usage similarity for a set of 10 nouns. The results show that our topic modelling approach outperforms the other two methods.

Read the paper · More papers on PaperTik