CLME: Robust Screen-Shooting Watermarking With Contrastive Learning and Mask-Guided Embedding
Jiaxing Liao, Jiaohua Qin, Yuanjing Luo, Wenyan Pan, Xuyu Xiang, Yun Tan · IEEE Transactions on Circuits and Systems for Video Technology · 2025
Screen-shooting watermarking technology plays a critical role in copyright protection and traceability. However, existing methods often lack sufficient robustness under strong noise interference and tend to introduce noticeable visual artifacts when embedding watermarks in smooth image regions, thereby degrading visual quality and increasing the risk of watermark exposure. To address these limitations, this paper proposes a Contrastive Learning and Mask-guided Embedding (CLME) framework for robust screen-shooting watermarking. The framework comprises two key components: (1) a mask-guided watermark embedding module that utilizes a Residual Dense Feature Extraction Block (RDFEB) and an Attention Mask Generation Block (AMGB) to adaptively embed watermarks into texture-rich regions, improving watermark invisibility; and (2) a contrastive learning-based watermark decoding network that employs contrastive loss to enhance the consistency of decoded features by treating features from the same watermarked image under different noise conditions as positive samples and features from different watermarked images as negative samples, thereby improving the robustness of watermark extraction. Experimental results demonstrate that the proposed CLME framework outperforms existing methods in terms of both robustness and visual quality. Specifically, at a shooting distance of 100 cm and a shooting angle of 40°, the watermark extraction accuracy reaches 99.58%, and the peak signal-to-noise ratio (PSNR) of the watermarked images reaches 42.624 dB, highlighting the framework’s strong potential for real-world applications.