Improving the Accuracy of Sarcasm Detection in Text Data Using a Smooth Support Vector Classification Model with Word-Emoji Embedding for News and Indian Indigenous Languages
N. Subalakshmi, R Babubalaji · 2024
By creating a Smooth Support Vector Classification (SSVC) model linked with word-emoji embedded data that is especially suited for newspapers and native Indian languages, the study aims to improve sarcasm detection in textual information. Because sarcasm is context-dependent and subtle, it poses a considerable barrier to processing natural speech, especially in multilingual situations. Word-emoji embedding is used by the proposed SSVC model to capture nuanced textual clues, emotions, and moods frequently overlooked by existing methods. The more emotional context that emojis offer to language data, the more precisely the algorithm can identify sarcasm. The paper tackles the complexity of multilingual sarcasm, concentrating on Indian indigenous languages, which are frequently underrepresented in NLP studies and lack adequate annotated information. The model performs better in classification than typical SVM models, as evidenced by extensive trials on real-world datasets of news items and social media chats. With applications in improving the interaction between humans and computers controlling online material, the study advances sentiment analysis, especially in various language circumstances.