Ensemble Learning-Based Sarcasm Detection in Hinglish Tweets Using Word2Vec Embedding
Archana Acharya, Rajeev Goyal · 2025
Sarcasm detection in social media content has gained significant attention due to its implications in natural language processing (NLP) tasks, such as sentiment analysis and conversational AI. This paper presents a methodology for detecting Hinglish sarcasm using machine learning (ML) algorithms with a focus on Word2Vec embedding for word representation and ensemble learning for improved classification performance. This paper utilized a dataset containing sarcastic tweets in Hinglish, a blend of Hindi and English, to train and evaluate various ML models, including Logistic Regression, Random Forest, Support Vector Classifier (SVC), Gradient Boosting, and Naive Bayes. The models were trained on tokenized and preprocessed text data, and their performances were assessed based on accuracy, precision, recall, and F1-score. Among the individual models, Gradient Boosting outperformed the others in terms of precision and F1-score. Additionally, we applied ensemble learning techniques— soft and hard voting classifiers—which demonstrated comparable performance to the best individual models, further enhancing the robustness of sarcasm detection. The results showed that ensemble models, particularly the hard voting classifier, achieved superior performance across all evaluation metrics. This study highlights the effectiveness of combining traditional machine learning algorithms with ensemble learning strategies in handling the challenges of sarcasm detection in code-mixed language data. Future work could explore the use of deep learning models and expand the dataset to improve accuracy and generalization to diverse contexts and languages.