Swim VQA: A Multimodal Approach to Drowning Detection by Fusing VQA and LiDAR

Muhammad Zeeshan Khan, Anuroop Gaddam, Dhananjay Thiruvady, Aaron Story · 2024

Visual Question Answering (VQA), a challenging field combining computer vision and natural language processing, is finding applications in critical real-world scenarios. This paper introduces a novel VQA approach to analyse swimmer behaviour and enhance drowning detection in swimming pools, leveraging the capabilities of LiDAR sensors. We present Swim-VQA, a new dataset comprising images of swimmers with corresponding textual questions and binary (yes/no) answers. This dataset, coupled with depth information from LiDAR, facilitates a deeper understanding of swimmer behaviour by enabling the analysis of visual, textual, and spatial information within the context of pool safety. We benchmark the Swim-VQA dataset using state-of-the-art Vision Language Models (VLMs), including Contrastive Language-Image Pre-Training (CLIP) and Bootstrapping Language-Image Pre-training (BLIP), as well as traditional Convolutional Neural Networks (CNNs) with Long Short Term Memory (LSTMs). Our results demonstrate the potential of VQA, particularly when enhanced with LiDAR data, for automated drowning detection and lay the groundwork for future research in this vital area.

Read the paper · More papers on PaperTik