Prototype-driven expert voting for pseudo-label selection with dual-balanced loss for imbalanced medical image classification
Sulei Wang, Sulei Wang, Shihao Sheng, Tiegong Wang, Tiegong Wang, Xudong Liang, Xiao Chen, Hao Shen, Haoyang Zhang, Jiang Xie, Guangchao Wang, Jiacan Su · Neurocomputing · 2026
Medical image datasets inherently suffer from class imbalance, label scarcity, and intra-class heterogeneity. Under semi-supervised learning, pseudo-label selection remains a central research focus. Class prototypes constructed in the latent space have been studied for pseudo-label selection, particularly because sub-class prototypes can effectively represent within-class heterogeneity in medical imaging datasets. However, since sub-class prototypes are constructed using a relatively small number of samples, they may lead to local optima during prototype updating and usage. In addition, loss weight rebalancing is commonly used to mitigate attention bias in model. Existing methods tend to focus exclusively on improving the minority class, without taking into account ambiguous samples, which are often difficult to learn and contain valuable information. Here, a prototype-driven framework PDMatch is proposed, incorporating Prototype-Driven Expert Voting (PDEV), Majority-class Pseudo-label Under-Sampling (MPUS) and Dual-Balanced Loss (DBL). First, an Expert Prototype Set is constructed, and hierarchical updating is proposed to sufficiently capture the feature distributions. PDEV then generates and filters pseudo-labels through prototype-driven voting, where multiple candidate sub-class prototypes are jointly considered. Second, MPUS offers a concise post-processing strategy that reduces imbalance in the pseudo-label distribution. Finally, Dual-Balanced Loss rebalances loss weights to jointly account for class distribution and sample difficulty. Comprehensive experiments are conducted on three datasets with distinct imbalance structures, demonstrating that our model addresses inherent data challenges and consistently outperforms state-of-the-art methods across multiple metrics. Additionally, we compare our model with clinical expert assessments to underscore its practical value in real-world applications. Source code is available at https://github.com/wang-su-lei/PDMatch .