Leveraging Multimodal Knowledge for Spatio-Temporal Action Localization

Keke Chen, T. Zhewei, Xiangbo Shu · 2024

Locating persons and recognizing their actions in chaotic scenes is a more challenging task in video understanding. Contrary to the quotidian conduct observed amongst humans, the actions exhibited during chaotic incidents markedly diverge in execution and the resultant influence on surrounding individuals, engendering a heightened complexity. Conventional spatio-temporal action localization methods that rely solely on a single visual modality fall short in complex scenarios. This paper explores STAL from a multimodal perspective by leveraging large language models (LLMs) and vision-language (VL) foundation models. We analyze the inherent feature aggregation phase of visual STAL and introduce a knowledge aggregation approach tailored for VL foundation models, termed Multimodal Foundation Knowledge Integration (MFKI). MFKI includes a generic decoder that facilitates the association of knowledge from VL foundation models with action features, as well as a specific decoder for relational reasoning in visual STAL. MFKI combines generic visual representations with specific video features to meet the demands of complex STAL tasks. Additionally, we utilize LLMs (i.e. GPT) and prompting to enrich label augmentation, fostering a more comprehensive linguistic understanding of complex actions. Experiments on the Chaotic World dataset have proven the effectiveness of the method proposed in this paper. The code is available at https://github.com/CKK-coder/Chaotic_World/tree/master.

Read the paper · More papers on PaperTik