Tweet2Vec: Character-Based Distributed Representations for Social Media
Bhuwan Dhingra, Zhong Zhou, Dylan Fitzpatrick, Michael Muehl, William W. Cohen · 2016
Text from social media provides a set of challenges that can cause traditional NLP approaches to fail.Informal language, spelling errors, abbreviations, and special characters are all commonplace in these posts, leading to a prohibitively large vocabulary size for word-level approaches.We propose a character composition model, tweet2vec, which finds vectorspace representations of whole tweets by learning complex, non-local dependencies in character sequences.The proposed model outperforms a word-level baseline at predicting user-annotated hashtags associated with the posts, doing significantly better when the input contains many outof-vocabulary words or unusual character sequences.Our tweet2vec encoder is publicly available 1 .