Collision attack and error corrected multimodal-inspired framework for underwater video enhancement

Jingchun Zhou, Chunjiang Liu, Dehuan Zhang, Zongxin He, Zifan Lin, Qiuping Jiang · Pattern Recognition · 2025

• Propose a multimodal-inspired underwater video enhancement framework. • Introduce collision-aware noise injection to simulate feature-level conflicts. • Design a hybrid ViT-convolutional architecture with low computational cost. • Achieve improved clarity and temporal consistency across underwater frames. Underwater video enhancement addresses degradation from absorption, scattering, and turbidity, which hinders visual tasks such as detection and tracking. Unlike traditional single-frame methods, we treat temporal cues, occlusion patterns, and structured degradations as implicit modalities within video streams. To this end, we propose CAECNet (Collision-Attack Error Correction Network), a novel multimodal-inspired underwater video enhancement network that integrates the ‘Collision Attack’ training strategy with the ‘Error Correction’ mechanism. This network overcomes the limitations of traditional single-frame enhancement methods, implementing high-precision and real-time inference. It improves temporal perception through multi-frame fusion and utilizes previous frames to assist real-time inference, meeting the demands of dynamic processing. By incorporating a Vision Transformer (ViT) and a lightweight depthwise separable convolution module, the network enhances spatial feature representation and computational efficiency. A branching-based error correction upsampler is designed to correct feature representation errors and reduce information entropy loss, thereby improving video detail restoration quality. The “Collision Attack” training strategy injects structured noise to accelerate network feature learning and reduce computational costs. Experimental results show that CAECNet significantly outperforms existing methods on multiple underwater video datasets, improving image clarity, inter-frame consistency, and computational efficiency, making it suitable for underwater robotic intelligent perception tasks.

Read the paper · More papers on PaperTik