Scene Retrieval in Traffic Videos with Contrastive Multimodal Learning

Touseef Sadiq, Christian Walter Peter Omlin · 2023

Retrieval of scenes from traffic videos is an important task in intelligent transportation systems (ITS) for efficient traffic management in AI smart cities. This work proposes natural language-based vehicle retrieval from traffic monitoring videos, emphasizing the significance of temporal information and context. We present contrastive learning as a technique to optimize joint representations of vision and language modalities within a shared latent representation space. The approach involves training contrastive losses to keep similar encodings closer in joint feature representation space by minimizing the distance between positive visual-text pairs and maximizing the distance between negative visual-text pairs. Our study employs state-of-the-art vision models for visual encoding and transformer-based language models for text encoding. We analyze the impact of feature selection from visuals and text on retrieval performance. We evaluate the efficacy of our proposed method on the AI City Challenge 2022 dataset for natural language-based vehicle retrieval, achieving performance accuracy of 49.84% Mean Reciprocal Rank (MRR) on the test dataset, securing second position on the leader board. Our approach highlights the effectiveness of feature selection and contrastive learning for enhancing multimodal retrieval tasks.

Read the paper · More papers on PaperTik