Diverse and styled image captioning using singular value decomposition‐based mixture of recurrent experts
Marzi Heidari, Mehdi Ghatee, Ahmad Nickabadi, Arash Pourhasan Nezhad · Concurrency and Computation Practice and Experience · 2022
Abstract With significant advances in vision and natural language processing, the generation of image captions becomes a need. Mathews, Xie, and He extended a new model to generate styled captions by separating semantics and style. In continuation of their work, here, a new captioning model is developed, including an image encoder to extract the features, a mixture of recurrent networks to embed the set of extracted features to a group of words, and a sentence generator that combines the obtained words as a stylized sentence. This Mixture of Recurrent Experts (MoRE) system uses a new training algorithm that derives singular value decomposition from weighting matrices of Recurrent Neural Networks (RNNs) to increase the diversity of captions. Each decomposition step depends on a distinctive factor based on the number of RNNs in MoRE. The used sentence generator gives a stylized language corpus without paired images. Besides, the styled and diverse captions are extracted without training on a densely labeled or styled dataset. MoRE on the COCO dataset generated diverse and stylized image captions without the necessity of extra‐labeling and improved descriptions in terms of content accuracy.