Scene Recognition and Narration from Video using Deep Learning Techniques

Rajesh Kumar Gunupudi, Sai Ramya Achanta, Dinesh Chandra D P, Hriday Rao Ganga, Niharika Vipparthy, Sreenivasa Rao Annaluri · 2021

The world becomes just a little harder if one of our senses are impaired and if that impaired sense is sight, every facet of our life becomes more challenging? Not every blind person can have a guide. This thought prompted us to work on a project called scene recognition and narration in which the application will recognize and understand the scene presented and dynamically give a verbal output. This application will help in the reduction of accidents and help make the lives of blind people just a little easier. The findings appear to provide good evidence that using widely accessible dense building blocks to approximate the predicted ideal sparse structure is a feasible technique for enhancing neural networks for computer vision. There are two more advantages to this architecture: One is that the number of units in each step is greatly increased, yet the computational complexity is kept under control. A significant number of input filters can be shielded from the final step to the next layer when dimensionality reduction is used widely. Second, visual data is processed on several scales and then aggregated so that the following phase may extract characteristics from many scales at the same time. YOLO Algorithm is used for the proposed architecture.

Read the paper · More papers on PaperTik