MBS: A Modality-Balanced Strategy for Multimodal Sample Selection
Yuntao Xu, Bing Chen, Feng Hu, Jiawei Liu, C. L. Zhao, Hongtao Wu · Machine Learning and Knowledge Extraction · 2026
With the rapid development of applications such as edge computing, the Internet of Things (IoT), and embodied intelligence, massive multimodal data are continuously generated on end devices in a streaming manner. To maintain model adaptability and robustness in dynamic environments, incremental learning has gradually become the core training paradigm on edge devices. However, edge devices are constrained by limited computational, storage, and communication resources, making it infeasible to retain and process all data samples over time. This necessitates efficient data selection strategies to reduce redundancy and improve training efficiency. Existing sample selection methods primarily focus on overall sample difficulty or gradient contribution, but they overlook the heterogeneity of multimodal data in terms of information content and discriminative power. This often leads to modality imbalance, causing the model to over-rely on a single modality and suffer performance degradation. To address this issue, this paper proposes a multimodal sample selection strategy based on the Modality Balance Score (MBS). The method computes confidence scores at the modality level for each sample and further quantifies the contribution differences across modalities. In the selection process, samples with balanced modality contributions are prioritized, thereby improving training efficiency while alleviating modality bias. Experiments conducted on two benchmark datasets, CREMA-D and AVE, demonstrate that compared with existing approaches, the MBS strategy achieves the most stable performance under medium-to-high selection ratios (0.25–0.4), yielding superior results in both accuracy and robustness. These findings validate the effectiveness of the proposed strategy in resource-constrained scenarios, providing both theoretical insights and practical guidance for multimodal sample selection in learning tasks.