Building a Semi-Supervised Dataset to Train Journalistic Relevance Detection Models
Nuno Guimar�ães, Álvaro Figueira · 2017
Annotated data is one of the most important components for supervised learning tasks. To ensure the reliability of the models, this data is usually labeled by several human annotators through volunteering or using Crowdsourcing platforms. However, such approaches are unfeasible (regarding time and cost) in datasets with an enormous number of entries, which in the specific case of journalistic relevance detection in social media posts, is necessary due to the wide scope of topics that can be considered relevant. Therefore, with the goal of building a relevance detection model, we propose an architecture to build a large scale annotated dataset regarding the journalistic relevance of Twitter posts (i.e. tweets). This methodology is based on the predictability of the content in Twitter accounts. Next, we used the retrieved dataset and build relevance detection models, combining text, entities, and sentiment features. Finally, we validated the best model through a smaller manually annotated dataset with posts from Facebook and Twitter. The F1-measure achieved in the validation dataset was 63% which is still far from excellent. However, given the characteristics of the validation data, these results are encouraging since 1) our model is not affected by content from other social networks and 2) our validation dataset was restrained to a specific time interval and specific keywords (which can affect the performance of the model).