Headlines data for social media popularity prediction

Alicja Piotrkowicz · University of Leeds · 2017

This dataset is part of a larger project on using headlines to predict the social media popularity of news articles. The dataset consists of two headlines corpora -- The Guardian and New York Times -- collected in 2014 using news outlet APIs. Each corpus includes a unique headline identifier (to enable recreating the corpus by querying the relevant API), the extracted features (news values, style, metadata), and the corresponding popularity on Twitter and Facebook.

Read the paper · More papers on PaperTik