Incomplete multimodal industrial anomaly detection via cross-modal distillation
Wenbo Sui, Daniel Lichau, Josselin Lefèvre, Harold Phelippeau · Information Fusion · 2025
Recent studies of multimodal industrial anomaly detection (IAD) based on 3D point clouds and RGB images have highlighted the importance of exploiting the redundancy and complementarity among modalities. However, achieving multimodal IAD in practical production lines remains a work in progress as it considers the trade-offs between the costs and benefits associated with the introduction of new modalities, while ensuring compatibility with current processes. Existing quality control processes combine rapid in-line inspections, such as optical and ultrasound imaging with high-resolution but time-consuming near-line characterization techniques, like industrial CT and electron microscopy to manually or semi-automatically locate and analyze defects in the production of Li-ion batteries and composites. Given the cost and time limitations, only a subset of the samples can be inspected by all methods, and the remaining samples are only evaluated through one or two forms of in-line inspection. To fully exploit data for deep learning-driven defect detection, the models must have the ability to leverage multimodal training and handle incomplete modalities during inference. In this paper, we propose CMDIAD , a Cross-Modal Distillation framework for IAD to demonstrate the feasibility of a Multi-modal Training, Few-modal Inference (MTFI) pipeline. Our findings show that the MTFI pipeline can more effectively utilize incomplete multimodal information compared to applying only a single modality for training and inference. Moreover, we investigate the reasons behind the asymmetric improvement using point clouds or RGB images for inference. This provides a foundation for our future multimodal dataset construction with additional modalities from manufacturing scenarios. The code is released on https://github.com/evenrose/CMDIAD . • A Cross-Modal Distillation framework is proposed for incomplete multimodal IAD. • A MTFI pipeline enables the model to be trained with multimodal data but inference with less. • The improvement of MTFI inference with point clouds is superior than that based on RGB images. • Associated information across modalities is crucial for completing missed information.