Modeling coherence and diversity for image paragraph captioning
Xiangheng He, Xinde Li · 2020
Image paragraph captioning task aims to provide detailed and objective multi-sentence descriptions of a target image. Existing models suffer from the problems of the monotony of sentence structure, the mechanical repetition of sentence meanings and stiff inter-sentence transitions in paragraph-level descriptions. In this paper, we consider modeling human concerned evaluation criteria, namely, diversity and coherence, and use them as explicit rewards in the proposed Reinforcement Learning based training strategy. Ablation study on the standard paragraph description dataset shows the necessity of leveraging Self-critical Sequence Training strategy and the effectiveness of our novel designed reward combined with CIDEr. Final experimental results show that our model improves the current state-of-the-art model in terms of four out of six standard evaluation metrics.