Learning a Contextualized Multimodal Embedding for Zero-shot Cooking Video Caption Generation

Lin Wang, Hongyi Zhang, Xingfu Wang, Yan Xiong · 2023

This paper proposes CookingCLIP, which introduces the latest CLIP (Contrastive Language-Image Pre-training) embedding from the general domain into the specific domain of cooking understanding, and makes two adaption upon the original CLIP embedding for better customization to the cooking understanding problems: 1) from the upstream perspective, we extend the static multi-modal CLIPembedding with a temporal dimension, to facilitate context-aware semantic understanding; 2) from the downstream perspective, we introduce the concept of zero-shot embedding to sequence-to-sequence dense prediction domains, facilitating CLIPbeing not only capable of telling “Which” (cross-modal recognition), but also capable of telling “When” (cross-context localization). Experiments conducted on two challenging cooking caption generation benchmarks, YouCook and CrossTask, demonstrate the effectiveness of the proposed embedding.

Read the paper · More papers on PaperTik