Weakly-supervised sentence-based aspect category and sentiment classification
Olaf Wallaart, Flavius Frăsincar, Finn van der Knaap · Knowledge-Based Systems · 2025
Sentiment analysis extracts the sentiment of content creators, enabling users to easily gain valuable insights from such data. Most existing methods rely on supervised learning approaches using labeled data. However, the retrieval of such labeled training data is difficult and expensive, especially for new domains and/or languages. This work focuses on simultaneously detecting aspect categories and sentiment polarities for a given sentence in a weakly-supervised setting. Two methods are proposed that combine an unsupervised labeling algorithm with a neural network architecture. The first proposed two-step model (SB-ASC) takes seed sentences as input for the labeling algorithm. By leveraging the power of pre-trained Sentence-BERT embeddings, the method is able to understand the contextual meaning of sentences to create a high-quality labeled dataset. This dataset is used by a class imbalance-robust BERT-based neural network that jointly learns latent features of aspect categories and the corresponding sentiment. The second proposed method (WB-ASC) uses the same neural network structure but takes seed words instead of seed sentences as input for the labeling algorithm. We conclude that SB-ASC outperforms WB-ASC as well as baselines and state-of-the-art weakly-supervised methods for aspect sentiment detection, achieving F1 scores for aspect category detection of 71.35%, 86.99%, and 73.86%, and F1 scores for sentiment classification of 89.24%, 89.98%, and 75.58% for the SemEval 2016 restaurant-5, restaurant-3, and laptop datasets, respectively. Furthermore, using domain-specific contextual language models boosts performance. • We predict aspect categories and sentiment polarities using weak supervision. • We propose two weak supervision methods. • We find that using seed sentences instead of words as input improves predictions. • We show that using domain-specific contextual language models boosts performance.