Joint Semantic Graph and Visual Image Retrieval Guided Video Copy Detection
Zhaohui Zhang, Xin Sheng Mao, Jiguang Zhang, Wenke Lian, Shibiao Xu, Xiaopeng Zhang · 2023
Recent years, video copy right infringement including secondary creation and video editing, have emerged one after another. The infringement examples of pirated videos are not limited to simple piracy or watermarking and other easily identifiable ways. The existing traditional methods mainly distinguish whether the video is copied by comparing the similarity of the video in the database. However, for video editing with large spatial and temporal differences, the cost of pixel-level comparison is huge, and the robustness of copy detection is also difficult to guarantee. Even the latest network learning methods still need to rely on enough training data and the generalization ability is poor. Aiming at the above issues, we propose a novel “coarse to fine” video copy accurate detection framework. In order to improve the detection efficiency of massive video frames, we first construct semantic graph for all key frames between source and target videos in the coarse detection stage. The pixel matching problem is transformed into graph structure matching to roughly divide the instances that may be copied, which significantly reduces the search range of subsequent video frames and improves the search speed. In the fine detection stage, we innovatively applied the cross-self-attention to video frame copy detection. By strengthening the contextual links between pixels in the keyframe space under the constraints of semantic mapping with the method of self-attention, global features are constructed by obtaining features common between the source video and the copy video through cross-attention to achieve the similarity measure between video frames in the high-dimensional space. Finally, based on the obtained global features, we perform an accurate and fast comparison through the hash segmentation method to achieve fast and high-precision video copy fine detection. Experimental results show that, compared with the existing state-of-arts our method has extremely high detection robustness against all mainstream video tampering results, and has the highest detection efficiency even in massive videos.