An Effective Machine Learning Approach with Hyper-parameter Tuning for Sentiment Analysis

Saima Kanwal, Ali Raza, Chunyan Bai, Dawei Zhang, Jing Wenn, Dileep Kumar · Data Intelligence · 2024

Sentiment analysis depends on individuals’ comments and opinions on events. Data from social media platforms like Twitter, Quora, or Facebook poses challenges due to informal language, including acronyms, misspellings, and ambiguous terms. Additionally, hyperparameters in machine learning models significantly impact performance. To address these issues, we propose advanced feature engineering techniques in Natural Language Processing (NLP) and hyperparameter optimization to enhance prediction accuracy and generalization capabilities. Our study employs Naïve Bayes, Logistic Regression (LR), Multi-layer Perceptron (MLP), and Support Vector Machine (SVM) to classify sentiments in tweets about Elon Musk’s potential acquisition of Twitter. The dataset, consisting of 100,000 tweets, is fetched using the Twitter representational state transfer application programming interface (REST API). We outline a sentiment analysis procedure to classify unstructured Twitter data, identify influential keywords, and categorize sentiments as Positive, Negative, or Neutral. Using a hybrid Lexicon NLP approach, we extract contextually significant emotionally charged words and assign sentiment polarities. Hyperparameter optimization via automated search methods ensures alignment with classifier performance estimates. SVM achieved an impressive accuracy rate of 97%. Cross-validation minimizes random variations, providing a reliable assessment of the model’s generalization capabilities, and demonstrating the method’s accuracy in predicting sentiments with larger new unseen standard datasets, and varying sentiment.

Read the paper · More papers on PaperTik