Transforming Videos to Text (VTT Task) Team: MMCUniAugsburg.
Philipp Harzig, Moritz Einfalt, Katja Ludwig, Rainer W. Lienhart · OPUS (Augsburg University) · 2020
The Multimedia and computer Vision Lab of the University of Augsburg participated in the VTT task only.We use the Auto-captions on GIF [1] (AC-GIF), MSR-VTT [2] and TRECVID-VTT [3] datasets for training our VTT models.We base our model on the Transformer [4] approach for both of our submitted runs, i.e., for runs 103.primary and 102.For our 103.primaryrun, we use the complete MSR-VTT dataset and 90% of the TRECVID-VTT dataset for pretraining while using the remaining 10% for validation.For the 102 run, we additionally utilize the complete AC-GIF dataset for the pretraining stage.Both runs were finetuned on TRECVID-VTT (90%).During finetuning the 102 run, the validation performance decreases significantly, while 103.primary improves in performance.The use of the AC-GIF dataset decreases the performance, because the domain and the captions are different to the other datasets.Overall, we find that training a Video-to-Text system on traditional Image Captioning pipelines [5] delivers very poor performance.When switching to a Transformer-based architecture our results greatly improve and the generated captions match better with the corresponding video (see Figure 3).