Extracting Top Trends from Twitter Discussions in Bulgarian

Boris Bankov · RePEc: Research Papers in Economics · 2017

Social networks offer plenty opportunities and areas for scientific research to dabble in user opinion mining and text analysis. The short text messages that get posted online present unique challenges related to automatic categorization and annotation. An interesting problem is the natural language filtering of text messages. Due to the huge volumes and sparsity of textual data machine learning algorithms are being applied. In this paper we take a look at the way to extract twitter messages in real-time containing Bulgarian texts. We also measure Twitter`s accuracy in terms of language identification from a 10 day dataset between 1st and 10th of October 2017. We propose a step by step text preprocessing algorithm, suitable for sanitizing tweets. We apply kmeans++ algorithm to cluster the extracted data and choose representative words for each cluster during each day.

Read the paper · More papers on PaperTik