Analysis of Social Media Data to Classify and Detect Frequent Issues Using Machine Learning Approach
Pankaj Bhowmik, Md. Sohrawordi, U. A. Md. Ehsan Ali, Md. Najmul Hasan, Prodip Kumar Roy · 2020
Social Media (SM) is becoming the next-level journalism as it reflects all the viral contents, social affairs, current conditions of our surroundings. SM site Facebook has achieved the utmost popularity in Bangladesh. Nowadays, people prefer to share their views about a problem on SM rather than any other platform. Considering these consequences, this study (Bangladesh perspective) proposed a system which will analyze Facebook data and will classify and detect the problems people are facing most frequently at a particular time. The problems are categorized into 12 major classes addressing the socio-economic aspects of Bangladesh. The dataset, contains public posts (and comments) from Facebook, is manually collected and labeled with the classes. Performing efficient data preparation, feature vectors are constructed with TF-IDF weights then Chi-square is applied to select the most correlated features. Four machine learning algorithms e.g., Logistic Regression, Support Vector Machine, Multinomial Naive Bayes, Random Forest are employed as supervised classifiers. Besides, hyperparameter tuning of these classifier algorithms is carried out with Grid Search and the best candidate model is selected applying 10-folds cross validation. Finally, the selected `Logistic Regression' model is built with securing classification accuracy of over 93%. The classification report illustrates important information about problems outgrowth.