Video Retrieval Model Based on Video Text Alignment

宇 张 · Journal of Image and Signal Processing · 2025

针对文本–视频检索遇到的全局对齐方法缺乏细粒度语义匹配以及跨模态语义鸿沟导致特征对齐困难的问题,提出一种高效全局–局部序列对齐方法(ETVA)。模型由文本编码器、视频编码器、文本–视频全局对齐模块和文本–视频细粒度对齐模块构成。其中文本编码器采用ALBERT模型,凭借其双向编码能力精准提取文本特征,能够提升跨模态特征的时序一致性与语义关联性。视频编码器利用多专家模块策略,从多模态、多特征角度全面捕捉视频信息。全局对齐模块通过聚合和变换特征,有效实现全局语义对齐;细粒度对齐模块基于共享聚类中心机制,深入挖掘文本和视频局部细节的语义关联。在实验中采用MSRVTT、ActivityNet Captions和LSMDC数据集,评价指标采用Recall@K和Median Rank,结果表明ETVA在不同数据集上均表现较好,在检索准确性相比其他方法有所提升。An efficient global local sequence alignment method (ETVA) is proposed to address the problem of global alignment methods lacking fine-grained semantic matching and cross modal semantic gaps leading to difficulty in feature alignment in text video retrieval. The model consists of a text encoder, a video encoder, a text video global alignment module, and a text video fine-grained alignment module. The text encoder adopts the ALBERT model, which accurately extracts text features with its bidirectional encoding ability, and can improve the temporal consistency and semantic correlation of cross modal features. The video encoder utilizes a multi expert module strategy to comprehensively capture video information from multiple modalities and feature perspectives. The global alignment module effectively achieves global semantic alignment by aggregating and transforming features; The fine-grained alignment module is based on a shared clustering center mechanism to deeply explore the semantic associations between local details in text and video. In the experiment, MSRVTT, ActiveNet Captions, and LSMDC datasets were used, and the evaluation indicators were Recall@K Compared with Median Rank, the results show that ETVA performs well on different datasets and has improved retrieval accuracy compared to other methods.

Read the paper · More papers on PaperTik