Learning Multiview Embeddings of Twitter Users
Adrian Benton, Raman Arora, Mark H. Dredze · 2016
Low-dimensional vector representations are widely used as stand-ins for the text of words, sentences, and entire documents.These embeddings are used to identify similar words or make predictions about documents.In this work, we consider embeddings for social media users and demonstrate that these can be used to identify users who behave similarly or to predict attributes of users.In order to capture information from all aspects of a user's online life, we take a multiview approach, applying a weighted variant of Generalized Canonical Correlation Analysis (GCCA) to a collection of over 100,000 Twitter users.We demonstrate the utility of these multiview embeddings on three downstream tasks: user engagement, friend selection, and demographic attribute prediction.