Image Based Action Recognition and Captioning Using Deep Learning

Ashutosh Singh, Sarishty Gupta · 2023

Image recognition and captioning are two essential tasks in computer vision that have gained significant attention in recent years. With the massive increase in digital images on the internet, there is a growing demand for automated systems that can accurately recognize and describe visual content. Image recognition involves classifying an image into predefined categories, whereas image captioning generates a natural language description of the image. These tasks require advanced machine learning algorithms and techniques to analyze and understand visual content, enabling computers to recognize objects, people, and scenes, and generate descriptions that are accurate and relevant. This paper focuses on using machine learning and deep learning algorithms for image recognition and captioning tasks. Specifically, the paper introduces a machine learning approach to image recognition using the SURF algorithm for feature extraction and SVM for classification, comparing the effectiveness of different kernels. The paper then explores the capabilities of deep learning in image captioning, using a pre-trained RESNET50 and VGG16 model for feature extraction and an LSTM-based neural network for generating captions. The paper also compares the effectiveness of the deep learning models by evaluating their accuracy using the BLEU score. The paper comprises two main components: an ML-based image recognition module and a DL-based image captioning module. The ML-based module provides robust image recognition capabilities, while the DL-based module generates high-quality image captions using a neural network-based approach. The paper highlights the importance of these tasks in various industries and concludes that leveraging these techniques can lead to innovative solutions in diverse domains.

Read the paper · More papers on PaperTik