Data Augmentation Using Multi-Turn Dialogue Prompting for Sentiment Analysis
Syifa Fatimah Azzahrah, Ade Romadhony · 2025
Label imbalance and data scarcity in Natural Language Processing (NLP) pose significant challenges to the development of effective text classification models. One approach to solve label imbalance and data scarcity is data augmentation. In this research, we examine the impact of multi-turn dialogue prompting approach on a large pretrained language model based chatbot for data augmentation on sentiment analysis task. Model evaluation on original dataset before data augmentation was performed shows accuracy of 0.6 and an average F1 score of 0.57. This performance reflects non-uniformity of labels and poor performance on the original dataset. After data augmentation was performed, the model performance improved with an accuracy score of 0.99 and F1-score of 0.99. In addition, data augmentation with single-turn dialogue is also performed. The model performance improved with an accuracy score of 0.92 and F-1 Score of 0.91. Although it can be able display satisfactory accuracy and F1 Score results, the model performance and data quality with data augmentation using multi-turn dialogue techniques are still much better than using single-turn dialogue techniques. The performance increase shows that data augmentation can significantly improve classification accuracy, especially for imbalance datasets.