YOLOv8 based Object Detection with Custom Dataset and Voice Command Integration
Ishita Singh, Nishtha Singh, Gracy Singh, Patteti Krishna · 2025
Assistive technologies for the visually impaired often lack efficiency, adaptability, and real- time responsiveness. This paper introduces a real-time object detection system based on YOLOv8, designed to process continuous video input, instantly identify objects, and convert detection results into speech. The system integrates a custom-trained YOLOv8 model for object recognition and a text-to-speech (TTS) engine for auditory feedback. Furthermore, voice command feature provides improved user interaction, allowing users to control the detection parameters. The proposed model has achieved a precision of 92.5%, recall of 88.3%, mAP@50 of 91.7%, and mAP@50-95 of 78.9% along with an inference speed of 16.5 milliseconds. Through model performance improvement via dataset augmentation, architectural improvements, and efficiency optimization, the proposed system achieves greater accuracy, faster processing speeds, and reduced computational complexity. These advancements contribute to the development of intelligent assistive systems that help contribute towards greater independence for visually impaired users.