Multi-label sentence classification using Bengali word embedding model
Md. Nowshad Hasan, Sourav Bhowmik, Md Mahfuzur Rahaman · 2017 3rd International Conference on Electrical Information and Communication Technology (EICT) · 2017
Multi-label sentence classification is a very popular technique to categorize text into several classes now-a-days. The classification process may vary based on the type of text used such as crime data, weather forecast data, accident data etc. We worked with only crime type news data from several famous newspapers. We made sub-categories such as victim, suspect, crime-type, crime-time, crime-place, police station, hospital and neutral for every sentences. We used two libraries LibSVM and Scikit-Learn to implement different kinds of algorithms like-SVM, Logistic Regression, Neural Network, Decision Tree, Linear Regression etc. We showed accuracy comparison using two libraries between different algorithms removing or having stop words in every sentence with 5000, 7500 and 10000 sentence corpus individually. We also showed the accuracy, F-score, recall and so on for every class we discussed before to observe how precisely our system can detect those classes. Our classifier can easily handle two or three classes for every sentence but most of the time it is not predictive for all of the classes of a sentence. We also discussed about how LibSVM and Scikit Learn work with multi-label classes in this paper.