Symmetrical Two-Stream with Selective Sampling for Diversifying Video Captions

Jin Wang, Yahong Han · 2024

Video captioning has garnered significant attention due to its promising potential in the field of cross-modal retrieval and analysis. However, the majority of existing methods tend to overlook the inconsistent quality of annotations in video-text pairs. Moreover, prevalent two-stream models tend to generate oversimplified captions, since these methods sacrifice the diversity feature in the original modality by learning the uniform representation of a shared subspace. In this paper, we introduce a weighting network for selective sampling on the training dataset, mitigating the interference of low-quality samples. Additionally, we adopt a symmetrical two-stream method to gain a better understanding of the shared cross-modal information for diversifying video captions. Experimental results on the MSVD and MSR-VTT have demonstrated excellent performance and effectiveness of our proposed method compared to existing methods across various metrics.

Read the paper · More papers on PaperTik