An Augmented Image Captioning Model: Incorporating Hierarchical Image Information
Nathan Funckes, Erin Carrier, Greg Wolffe · 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA) · 2021
Despite published accessibility standards many websites remain nan-compliant, containing images lacking accompanying textual descriptions. This leaves visually-impaired individuals unable to fully enjoy the rich wonders of the web. To help address this inequity, our research seeks to improve the ability of autonomous systems to generate accurate, relevant image descriptions. Our model enhances training efficacy by incorporating the use of category labels, high-level object superclasses, which are derivable using modern object-detection models. We show that this simple augmentation to an existing architecture results in a statistically significant improvement in caption quality.