Headlines data for social media popularity prediction
Alicja Piotrkowicz · University of Leeds · 2017
This dataset is part of a larger project on using headlines to predict the social media popularity of news articles. The dataset consists of two headlines corpora -- The Guardian and New York Times -- collected in 2014 using news outlet APIs. Each corpus includes a unique headline identifier (to enable recreating the corpus by querying the relevant API), the extracted features (news values, style, metadata), and the corresponding popularity on Twitter and Facebook.