Unlocking Cognitive Insights: Leveraging Transfer Learning for Caption Generation in Signed Language Videos

Ajay M. Pol, Shrinivas A Patil · 2024

The research delves into sign language recognition from video, employing a transfer learning approach. Established models like VGGNets, ResNets, DenseNet, Inception, and MobileNet are utilized to extract features from sequences of video frames. This enables the extraction of both local and global features, crucial for evaluating sign language performance. The final model utilizes an LSTM-based architecture, incorporating input from these feature sets alongside target text vectors. The analysis focuses on gesture-oriented sign language for word recognition, emphasizing the BLEU parameter. Through meticulous examination, the study pinpoints the ResNet-101 model as the most adept for the application, offering a comprehensive comprehension of sign language recognition within video contexts. This holistic approach contributes significantly to advancing the field's understanding and application, paving the way for more effective sign language recognition systems.

Read the paper · More papers on PaperTik