Fine-Tuning Fasttext Using Bayesian Optimization for Movie Review Sentiment Analysis

Rico Halim, Abba Suganda Girsang, Mohammad Faisal Riftiarrasyid · 2023

The advancements in technology during the 20th century resulted in the onset of the digital computer era. This study investigates the relative importance of earlier algorithms in comparison to more recent ones. Various research projects suggest improving vectorization by combining conventional and modern techniques, while others suggest optimizing it through algorithmic methods. This study primarily focuses on employing Bayesian optimization to optimize hyperparameters, hence improving the performance evaluation of TF-IDF FastText sentiment classification models. This study presented four models: the initial model utilized FastText with a preprocessing step that involved removing stopwords, the second model employed Fast-Text without any optimization using Bayesian optimization, the third model utilized FastText with optimization using Bayesian optimization, and the fourth model combined FastText with TF-IDF and was further optimized using Bayesian optimization. The Support Vector Machine (SVM) technique will be employed to evaluate all of the models. The findings suggest that the model’s performance stays unchanged when stopwords are eliminated (Precision, Recall, F1: 0.9007). The model demonstrates a slight enhancement through the utilization of Bayesian optimization, resulting in a 0.02% increase, reaching a final accuracy of 0.9019%. Bayesian optimization is frequently employed to improve performance. The performance of Bayesian optimization is affected by the magnitude of the learning rate, and models that do not utilize Bayesian optimization demonstrate inferior performance. Both the TF-IDF FasText model and FastText model yielded comparable outcomes, attaining an F1 score of 0.9019 after being tuned by Bayesian optimization.

Read the paper · More papers on PaperTik