Design of a Multimodal Network Data Security Detection Model Based on Self‐Supervised Learning

WenCheng Lin · Security and Privacy · 2025

ABSTRACT As network attacks evolve in complexity, traditional methods for detecting network data security are increasingly inadequate, especially in handling multi‐modal data and addressing sample imbalance issues. To overcome these challenges, this paper introduces TransRL, a model for multi‐modal network data security detection built on self‐supervised learning. By integrating self‐supervised learning, Transformer, and reinforcement learning techniques, TransRL effectively extracts deep features and adapts to the dynamic nature of multi‐modal data. Experimental results on the CICIDS 2018 and NSL‐KDD datasets show that TransRL significantly outperforms conventional methods in key performance metrics such as accuracy, recall, and F1‐score. Notably, it achieves an accuracy of 96.8% and an F1‐score of 0.964 on the CICIDS 2018 dataset, and 95.4% accuracy and 0.941 F1‐score on the NSL‐KDD dataset. Additionally, ablation experiments highlight the importance of each module: the self‐supervised learning module enhances the model's handling of sparse samples, the Transformer module improves multi‐modal feature representation, and the reinforcement learning module refines real‐time detection. This work offers an effective solution for data security detection in complex network environments and provides useful insights for advancing multi‐modal learning and network security technologies.

Read the paper · More papers on PaperTik