Soccer captioning: dataset, transformer-based model, and triple-level evaluation
Ahmad Hammoudeh, Bastien Vanderplaetse, Stéphane Dupont · Procedia Computer Science · 2022
This work aims at generating captions for soccer videos using deep learning. The paper introduces a novel dataset, model, and triple-level evaluation. The dataset consists of 22k caption-clip pairs and three visual features (images, optical flow, inpainting) for 500 hours of SoccerNet videos. The model is divided into three parts: a transformer learns language, ConvNets learn vision, and a fusion of linguistic and visual features generates captions. The suggested evaluation criterion of captioning models covers three levels: syntax (the commonly used evaluation metrics such as BLEU-score and CIDEr), semantics (the quality of descriptions for a domain expert), and corpus (the diversity of generated captions). The paper shows that the diversity of generated captions has improved (from 0.07 reaching 0.18) with semantics-related losses that prioritize selected words. Semantics-related losses and the utilization of more visual features (optical flow, inpainting) improved the normalized captioning score by 27%.