Automated Image Captioning Systems
Aman Swaraj, R. Rubin Silas Raj, Aniket Shaw, Richa Dubey, Prianka Dey, Sagarika Chowdhury · 2025
The objective of the project is to design a system that leverages live camera feeds to detect objects in real time and generate concise, meaningful descriptions of the scene. The system processes the visual input, identifies key objects or actions, and produces a one-line caption summarizing the observed scenario. This caption is then saved in a text file, stored in the same directory as the program. For instance, if the live feed shows a person seated in front of the camera holding flowers, the system would generate the caption: “A person holding a bouquet of flowers,” which gets automatically written to the file. This solution integrates advanced technologies like real-time object detection and natural language generation to achieve efficient scene analysis and automated captioning. Such a system has diverse applications, including aiding visually impaired individuals, enhancing video surveillance, and improving user interaction in smart environments. By combining visual recognition with descriptive language processing, it provides an innovative way to summarize live camera feeds in a user-friendly format.