Enhancing Sentiment Analysis in Multilingual Social Media Data Using Transformer-Based NLP Models A Synthetic Computational Study

GOPICHAND BANDARUPALLI · 2025

Social media platforms host a diverse array of multilingual content, blending languages, slang, and emojis, which poses significant challenges for traditional sentiment analysis tools. This study evaluates transformerbased NLP, specifically multilingual BERT (mBERT), for sentiment classification across English, Hindi, and Spanish. Using a synthetic dataset of 10,000 postsdesigned to emulate real-world social media with codeswitching and informal expressions-we compare mBERT against logistic regression with TF-IDF and LSTM networks. Evaluated via accuracy (precision), recall, and ROC curves, mBERT achieved a peak accuracy of 0.91, outperforming baselines by 14-18%. This computational study demonstrates transformers' efficacy in decoding multilingual sentiments and identifies optimization opportunities for resource-efficient deployment.

Read the paper · More papers on PaperTik