Baselines for demographic inference on a new gold standard twitter corpus

Jason Pilny Radford, Luke Horgan, David M. J. Lazer · 2017

A variety of studies have shown that machine learning methods like convolutional neural nets and random forests can be used to accurately infer characteristics of people online such as their gender, age, race, or political orientation. However, these studies are based on labels generated using the data themselves, typically human coding of subjects, and presume subjects are authentic humans. This creates systematic selection biases owing what features humans can draw inferences from. In this preliminary study, we connect Twitter Data to an exogenous data source, public voter data, to create a new gold standard data set for inferring demographic information about online participants. We run a standard battery of machine learning algorithms on bag-of-words representations of individuals' twitter posts to generate new baselines for how well these characteristics can be predicted. Our baselines are substantially lower than most reported studies, suggesting sampling bias has led to an over-estimation of how well machine learning algorithms perform on this task.

Read the paper · More papers on PaperTik