Exploring Homophobic Discourse in Tamil Transcripts: A Comparative Analysis of Sentiment Classification Models using TF-IDF Vectorization
Soubraylu Sivakumar, Jyothi Prakash, Prathik Saravanan, T. Rajesh Kumar, Selvanayaki Kolandapalayam, R. Rajalakshmi · 2024
This research delves into the linguistic patterns of homophobic discourse within Tamil transcripts, employing sentiment analysis methods and TF-IDF vectorization for feature representation. Homophobia, characterized by prejudice, discrimination, or antagonism directed against individuals based on their sexual orientation, is a pervasive issue addressed in this study. Various ML algorithms, such as Support Vector Machine Classifier (SVM), Naive Bayes Classifier (NBC), Logistic Regression, and Gradient Boosting (GB), are utilized to discern and assess their effectiveness in distinguishing between homophobic and non-homophobic language. The dataset has been derived from various online platforms, prominently including YouTube, alongside other sources. Performance metrics including accuracy, F1 score, precision, and support are employed to gauge the models' performance. The outcomes shed light on the prevalent discriminatory language within the Tamil-speaking community and offer insights into combating homophobia through computational linguistic techniques. Notably, SVM demonstrated the highest accuracy of 88.18% in the classification task.