Image and Encoded Text Fusion for Deep Multi-Modal Clustering

Truong Dong, Kyungbaek Kim, Hyuk-Ro Park, Hyung-Jeong Yang · 2020

Deep multimodal clustering is challenging since it needs to learn appropriate features for different modalities and find correct clusters by using consistency among the modalities such as textual and visual domains. In this paper, we proposed a fused deep multimodal clustering (FDMMC) approach that fuses image and associated text description into unified information enriched image. To cluster feature representations of resulting images in complex space, the data was compressed into a lower-dimensional by the convolutional autoencoders. Then, iteratively minimizes the Kullback-Leibler divergence between the distribution of the encoded data and the target distribution. To prevent feature space from being distorted by the clustering loss, we keep the decoders remained to preserve the data structure. By integrating the clustering and autoencoder's reconstruction loss, FDMMC can simultaneously optimize cluster labels assignment and features refinement. The experimental evaluations on Coco-cross datasets illustrated the proposed method has better performance and the effectiveness of our algorithm in the image-text pairs clustering task.

Read the paper · More papers on PaperTik