Video Question Answering with Phrases via Semantic Roles

Arka Sadhu, Kan Chen, Ram Nevatia · 2021

Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases.These metrics limit the VidQA models' application scenario.In this work, we leverage semantic roles derived from video descriptions to mask out certain phrases, to introduce VidQAP which poses VidQA as a fillin-the-phrase task.To enable evaluation of answer phrases, we compute the relative improvement of the predicted answer compared to an empty string.To reduce the influence of language-bias in VidQA datasets, we retrieve a video having a different answer for the same question.To facilitate research, we construct ActivityNet-SRL-QA and Charades-SRL-QA and benchmark them by extending three vision-language models.We perform extensive analysis and ablative studies to guide future work.Code and data are public.Video description: A man on top of a building throws a bowling ball towards the pins Q4: throws a bowling ball towards the pins.Model's generated answer: A man standing on a house Correct answer: A man on top of a building Q5: A man on top of a building a bowling ball towards the pins.Model's generated answer: throws Correct answer: throws Q6: A man on top of a building throws towards the pins.Model's generated answer: a ball Correct answer: a bowling ball Q7: A man on top of a building throws a bowling ball Model's generated answer: towards some bottles Correct answer: towards the pins (b) Free-form Answer Generation ARG0

Read the paper · More papers on PaperTik