Deep visual embedding for image classification
Adel Saleh, Mohamed Abdel‐Nasser, Md. Mostafa Kamal Sarker, Vivek Kumar Singh, Saddam Abdulwahab, Nasibeh Saffari, Miguel Ángel García, Domènec Puig · 2018
This paper proposes a new visual embedding method for image classification. It goes further in the analogy with textual data and allows us to read visual sentences in a certain order as in the case of text. The proposed method considers the spatial relations between visual words. It uses a very popular text analysis method called ‘word2vec’. In this method, we learn visual dictionaries based on filters of convolution layers of the convolutional neural network (CNN), which is used to capture the visual context of images. We employee visual embedding to convert words to real vectors. We evaluate many designs of dictionary building methods. To assess the performance of the proposed method, we used CIFAR10 and MNIST datasets. The experimental results show that the proposed visual embedding method outperforms the performance of several image classification methods. Experiments also show that our method can improve image classification regardless the structure of the CNN.