Context-Aware Automated System for Image Caption to ASL Translation for Improved Accessibility in Media Applications
Oluwafolahanmi Aluko, Vineetha Menon · 2025
With the emergence of Artificial Intelligence (AI) and the rapid growth of Generative AI, significant progress has been made in image captioning, delivering remarkably precise captions from static images and, more recently, videos [1]. Alt-Text has significantly progressed due to initiatives aimed at improving the context and sentiment recorded during automated captioning. This paper seeks to elaborate on a portion concerning Natural Language Processing (NLP) and its significance in enhancing accessibility and context-awareness within the framework of the American Sign Language (ASL) application. This work presents a case study on significance of providing complete context regarding a caption or image to the end user, encompassing emotions and conscious attributes present within the phrase. This paper extensively elaborates on the accessibility functions presently involved in image-captioning. Is user-dependent (requiring users to supply a caption for the image) or standard machine-created alt-text sufficient for complete accessibility? Might incorporating a sign language, such as American Sign Language (ASL), be essential in providing users with more information than what the text of a caption communicates? The key question is: how can we enhance this image-caption translation to be more context-sensitive to convey the subtleties that could be crucial for user comprehension? The aim is to create a world where all voices resonate and every person feels accepted, and this paper represents an initial move towards an AI revolution that might be crucial in realizing that vision.