Cognitive Object Detection: A Deep Learning Approach with Auditory Feedback
S Pooja, Prasanna Kumar K, A S Sushmitha Urs, Vaibhavi B Raj, B R Madhu, Vinod Kumar S · 2024
Amidst the backdrop of technological convergence, this study delves into the realm of augmented visual intelligence. It does so by orchestrating a harmonious fusion of the YOLOv3 framework, deep learning paradigms, and the seamless integration of Text-to-Speech (TTS) capabilities. The research paper meticulously dissects a methodological blueprint designed for the creation of an object recognition system, impeccably tailored to the precise identification of objects within the comprehensive COCO dataset. With the YOLOv3 architecture as our cornerstone, meticulous parameter fine-tuning transpires within the Darknet framework, ensuring an unswerving alignment with the diverse object categories that define the COCO dataset. Our system demonstrates proficiency by deploying TTS technology to deliver real-time auditory interpretations of recognized objects, enhancing both user accessibility and engagement. The ethical compass steadfastly guides our approach, encompassing privacy safeguards that underscore our commitment to the conscientious and responsible utilization of data. System performance is rigorously assessed through the lens of pivotal metrics, including precision, recall, and the F1 score, validating the system’s precision and reliability. This research elucidates the transformative potential innate to the amalgamation of deep learning and TTS integration within the sphere of object recognition, thus carving a path for pioneering applications and the evolution of technology.