Learning a Video-Text Joint Embedding using Korean Tagged Movie Clips
Gyeong-June Hahm, Chang-Uk Kwak, Sun-Joong Kim · 2020
For intelligent multimedia services, video contents understanding is a major challenge. In the existing video retrieval approaches, manual descriptive sentence data is necessary for retrieving desired videos against user's search intent. To overcome these limitations, modeling visual concepts included in video and sentence is necessary to learn a mapping of video and text into a common vector space, where relevant videos and texts are close to each other. In this study, we construct a new dataset containing 250 Korean movies with manual text description in Korean. Also, video-text joint embedding model and its quantitative and qualitative search results are introduced. With our proposed model, video manual tagging is no longer necessary for video retrieval services.