Classification of Malayalam-English Mix-Code Comments using Current State of Art
Subramaniam Kazhuparambil, Abhishek Kaushik · 2020 IEEE International Conference for Innovation in Technology (INOCON) · 2020
Automatic classification of YouTube comments is a challenge especially when the comments are multilingual, as the messages are often rife with slang, symbols and abbreviations of the respective vernacular. In this work, we have evaluated top-performing classification models for classifying comments which are a mix of English and Malayalam. The statistical analysis of results indicates that XLM was the top-performing model with an accuracy of 67.31 %. Multi-layer Perceptron (MLP) with Term Frequency vectorizer produced the best results out of all the Deep Learning models with an accuracy of 65.98%. Random Forest with Term Frequency vectorizer was the top-performing model out of all the traditional classification models with an accuracy of 63.59%.