Employing Tensor Decomposition and Contextual Inference for Reducing the Semantic Discrepancy in Image Processing
K. Suneetha, Namit Gupta · 2023
Despite the growing need for effective image tagging in both commercial and research domains, the inconsistency between image content and text annotations remains a significant issue. We propose an innovative approach that leverages deep learning and tensor decomposition to enhance the quality of image labelling and eliminate the reliance on labor-intensive manual tagging. Our methodology transforms images into tensors, facilitating the formation of a standardized feature space. We employ a sophisticated five-layer neural network architecture that integrates three-level tensor decomposition to match images with their context groups accurately. This structure is adept at capturing the nuanced patterns in data, which is crucial for reducing the semantic discrepancy between visual and textual information. We have rigorously evaluated our algorithm using the Corel-10K dataset, focusing on the specific case of penguin images, to demonstrate the efficacy of our model. The results showcase a marked improvement in the relevance of generated tags, thereby enhancing the retrieval accuracy and mitigating the issue of inappropriate tag assignment. This work not only contributes to the fields of Machine Learning and Data Mining by refining feature extraction and context estimation techniques but also has significant implications for the development of NLP algorithms that can better interpret and generate descriptive tags for visual content. Moreover, our approach sets the stage for future explorations into the integration of AI-driven image tagging systems with robotic systems for autonomous data handling and organization.