Enhancing Object Detection through Auditory-Visual Fusion on Raspberry Pi and FogBus
R. Raja Subramanian, Lagisetty Ravikiran, Kota Venkata Pavan Teja, Kondamarri Venkatesh Reddy, Kondeti Akarsh Chowdary · 2024
This study introduces a new method for improving object detection using audio-visual fusion, which is implemented on a Raspberry Pi and integrates with FogBus. In order to identify and characterize objects in real time, the system integrates auditory signals with visual data obtained by a camera. TensorFlow Lite is used to recognize objects, and FogBus controls processing and resource allocation via fog nodes to guarantee low latency and effective computation. Through the use of multimodal data fusion techniques, visual and auditory information are combined to provide visually impaired users with audio outputs that are translated into detailed scene descriptions. By utilizing edge and fog computing, the system maximizes performance while achieving accurate object detection. A comparative examination reveals that this model works better in real-world circumstances than conventional YOLO-based models, providing increased accuracy and efficiency. The study emphasizes how merging advanced computing frameworks with lightweight models can help develop scalable assistive technology.