Trans-Convo-Former Net for Hierarchical Prediction of Household Images
Divya Arora Bhayana, Om Prakash Verma · ACM Transactions on Multimedia Computing Communications and Applications · 2025
Image classification has become the backbone of computer vision in recent times. Hierarchical image classification has been a scarcely exploited field, particularly in household images. Although many convolution and transformer learning models have been introduced for image classification, the fusion models exhibit much better performance in image classification. The potential of hierarchical image classification for household robotics has not yet been explored. Therefore, we propose a Trans-convo-former net for the hierarchical prediction of household images. The fine class refers to the class of the object identified, and the coarse class refers to the object’s location. This process facilitates the path-planning stage of household robotics. Trans-convo-former net employs self-attention-based encoders with intermittent convolution layers to extract global and local features from the image. The model is observed to outperform the state-of-the-art models applied for hierarchical image classification. Trans-convo-former net is proposed in two versions namely; big and small. The model is compared to another fusion model as well. The performance of the proposed model is found to be the most optimum. An ablation study is also performed with different numbers of transconvoformers, attention layers, and epochs to find the best-performing parameters for the proposed model.