Clothing retrieval based by image and text combination
Zongbao Liang, Yang Liu, YunFei Yuan, Bo Chen, Feifei Tang · 2021
Image text combination retrieval is a new direction in multimodal retrieval, in which query is composed of image and modified text. The retrieved target image should not only be similar to the query image, but also have the change specified by the modified text. The traditional clothing retrieval adopts the single-mode retrieval method of image search or text search, which is lack of retrieval flexibility. To solve the problem of feature fusion caused by semantic differences between clothing image and text, this paper proposes a multi-dimensional feature fusion model, which constructs a high-dimensional visual feature and semantic feature fusion model based on the scaling point product attention mechanism to extract high-dimensional fusion features, then the low dimension visual semantic fusion features are used as the residual of high dimension fusion features for target image retrieval. Compared with the previous feature fusion methods, the recall rate of Top1 on Fashion200k data set is increased by 15.4%, which is obviously superior to most of the existing graph and text feature fusion models, which shows that the model is advanced and effective.