Multimodal Cascaded Framework with Multimodal Latent Loss Functions Robust to Missing Modalities
Vijay John, Yasutomo Kawanishi · ACM Transactions on Multimedia Computing Communications and Applications · 2025
Despite interest in multimodal classification, few studies have addressed the missing modality problem in which an incomplete multimodal input with one or more missing modalities is classified as the target class. The missing modality problem is shown to reduce the classification accuracy as the discriminative power of the obtained feature space is reduced. In this study, we address the missing modality problem in multimodal classification using a novel cascaded framework. The proposed framework is formulated in the feature space to address the missing modality problem by generating complete multimodal data from incomplete multimodal data. Subsequently, an optimal multimodal data is obtained by feature selection of the generated and original data. The proposed cascaded framework consists of three steps: feature extraction, feature generation, and classification. The framework is formulated to handle both complete and incomplete multimodal data simultaneously. The cascaded framework is trained using novel latent loss functions: missing modality joint loss, centroid joint loss, and latent prior loss. These loss functions, based on metric learning, are designed to ensure that data from the same class remain proximate in the latent space irrespective of the presence or absence of modality data. The cascaded framework is validated on bimodal audio-visible RAVDESS and trimodal audio-visible-thermal Speaking Faces datasets. The experimental results show that the cascaded framework improves classification accuracy even with incomplete multimodal data.