Understanding and Mitigating the Soft Error of Contrastive Language-Image Pre-training Models

Yihao Shi, Bo Wang, Shengbai Luo, Qingshan Xue, Xueyi Zhang, Sheng Ma · 2024

In recent years, MultiModal Large Language Models (MM-LLMs), based on the Contrastive Language-Image Pretraining models (CLIP), have achieved the best results in many fields. CLIP breaks through the gaps between language models and image models, realizes zero-shot image classification, and achieves excellent performance in tasks such as text-to-image generation, image style transformation, and long video generation. However, there are few studies on the fault tolerance of CLIP with soft errors, which hinders the application of multimodal large models in the field of security. Based on the analysis of the fault tolerance of common multimodal large models, we proposes a soft error mitigation framework. According to the experiments in this paper, the framework can effectively detect soft errors and mitigate the errors.

Read the paper · More papers on PaperTik