Prediction of Actions and Objects through Video Analysis Using Stepwise Prompt
Tsukasa Hirano, Kengo Ozaki, Takeshi Morita · 2024
Against the backdrop of the aging of society in Japan and the frequent occurrence of accidents involving elderly people in ”homes,” we took on the 2nd International Knowledge Graph Challenge, whose goal is ”developing an AI system that detects accident risks latent in daily life in the home with concrete explanations and presents safer alternatives. The task of this challenge is to ”answer questions on a multimodal dataset consisting of videos and knowledge graphs that represent daily activities.”. The task consists of two subtasks: Task 1 answers questions by searching knowledge graphs, and Task 2 answers questions that need help to be answered alone on knowledge graphs with missing data. In Task 1, we answered four of eight questions by combining SPARQL query searching and Python processing. In Task 2, we conducted three validation experiments using existing Vision-Language Models, OpenAI CLIP, OpenAI GPT-4V, and LLaVA, with the policy of filling in missing behaviors and objects from videos. We conclude that LLaVA is suitable for predicting missing data. The source code used in the experiment is available at https://github.com/H-Tsukasa/kg_challenge_team1.