Caption Prompt for Video Moment Retrieval and Highlight Detection

Qing-Hua Ling, Zicheng Liu, Haifeng Zhao · 2025

The aim of video moment retrieval and highlight detection is to locate moments in videos and estimate the saliency scores of video clips given user queries. Although models based on audio information have made some progress, they are unable to effectively utilize the audio information in videos, and it is difficult to achieve interaction between cross-modal features. To address these issues, our design a new Transformer-based model called Caption Prompt based DEtection TRansformer (CP-DETR). The main idea of this model is to extract the most effective caption information from the audio, calculate the similarity within a single modality using text queries and captions as prior information for the video moment retrieval and highlight detection tasks, and guide the completion of these tasks. With the introduction of audio information, we observe a significant improvement in model performance on the three benchmarks of QVHighlights, Charades-STA, and TVSum, demonstrating the effectiveness of the proposed method.

Read the paper · More papers on PaperTik