A Novel Transformer-based Two Stage Framework for Multi-Label Image Classification
Zhijie Wang, Naikang Zhong, Xiao Lin · 2024
Multi-label image classification (MLIC) has received much attention due to its wide applications. Recently, the Transformer-based methods have exhibited excellent performance on MLIC tasks. Existing Transformer-based methods either used standard encoders that employ traditional relative positional encoding, or discarded encoders by focusing on the decoders to establish relations between labels and the regions of interest (RoI) in images. Yet, they pay less attention on enhancing the encoders. Moreover, the self-attention layer in the decoder is often employed in existing works, and is considered a component that could potentially enhance the internal relationships of label embeddings. It yet is unclear whether the self-attention layer is really helpful in establishing label correlation. Last but not least, existing Transformer-based methods use Asymmetric Loss to alleviate the sample imbalance problem; however, Asymmetric Loss may lead to the exclusion of some negative samples that are actually helpful for model training. To address these issues, this paper presents a novel Transformer-based two stage framework, which can be viewed as a fusion of the RoI based technique and an adapted Transformer. Our framework captures global and local features in model training. It is simple and easy-to-implement, but can achieve excellent performance. We have conducted extensive experiments on two widely used public datasets. The results consistently show us that our proposed method is feasible and also competitive, compared against state-of-the-art models.