Textual emotion detection with complementary BERT transformers in a Condorcet’s Jury theorem assembly
Gerardo Bárcena Ruiz, Richard de Jesús Gil Herrera · Knowledge-Based Systems · 2025
• Novel approach applying Condorcet’s Jury Theorem (CJT) to NLP for Textual Emotion Detection (TED) using BERT-based transformers . The study suggests that aggregating multiple independent and competent classifiers can lead to more accurate decision-making, which is leveraged here for natural language processing (NLP) tasks. • Jury-Based Ensemble Voting Mechanism . Each transformer acts as an independent juror, classifying text, but a majority voting scheme that follows the principles of CJT produces the final classification. The ensemble approach aims to improve prediction robustness by leveraging model diversity. • Significant Performance Improvement . The proposed algorithm achieved an F1-Score of 0.958, notably outperforming single models, which averaged 0.927. These F1-Score values confirm that aggregating multiple independent models outperforms individual models. This paper explores a novel approach to textual emotion detection (TED) in Spanish and English, leveraging an ensemble of partially trained BERT transformers within a Condorcet’s Jury Theorem (CJT) framework. Recognizing the challenges of limited training data and the complexities of emotion classification, this research investigates whether a combination of BERT models in the CJT ensemble can enhance performance even when individual models have incomplete training. The study evaluates different BERT modalities (BERT, RoBERTa, DistilBERT) and datasets, including SemEval-2018, XED, and Dair-ai/emotion. The main contribution is the development of a CJT ensemble, specifically the Jury Dynamic (JD), a key contribution of this research. This algorithm is designed for deployment in unsupervised production environments, eliminating the need for labeled data or continuous human supervision, leveraging Reinforcement Learning (RL). The Jury Dynamic (JD) adapts to incoming data, making it suitable for real-time applications. Experiments involve retraining BERT models with varying levels of emotional data reduction to simulate incomplete training. Results demonstrate that the CJT ensemble, particularly the JD, can effectively mitigate the negative impacts of limited training data, achieving comparable performance to fully trained models and outperforming individual models. The study highlights the importance of high-quality datasets for TED, particularly in Spanish, and proposes future research directions, including the evaluation of various classifiers and ensemble configurations.