Detecting Newsworthy Topics in Twitter
Steven Van Canneyt, Matthias Feys, Steven Schockaert, Thomas R. DeMeester, Chris Develder, Bart Dhoedt · Ghent University Academic Bibliography (Ghent University) · 2014
The task of the SNOW 2014 Data Challenge is to mine Twitter streams to provide journalists a set of headlines and complementary information that summarize the most newsworthy topics for a number of given time intervals.We propose a 4-step approach to solve this.First, a classifier is trained to determine whether a Twitter user is likely to post tweets about newsworthy stories.Second, tweets posted by these users during the time interval of interest are clustered into topics.For this clustering, the cosine similarity between a boosted tf-idf representation of the tweets is used.Third, we use a classifier to estimate the confidence that the obtained topics are newsworthy.Finally, for each obtained newsworthy topic, a descriptive headline is generated together with relevant keywords, tweets and pictures.Experimental results show the effectiveness of the proposed methodology.