UTS ISA Submission at the TRECVID 2019 Video to Text Description Task.
Qi Rao, Guang Li, Yi Ping Yang, Feng Zhang, Ziwei Wang · UTS ePRESS (University of Technology Sydney) · 2019
In this paper, we summarize the technical details applied in our submission of TRECVID 2019[1] video to text task. The main effective improvements include three parts: Several efficient and comprehensive high-level features to gain expressive visual feature encodings, the algorithms in regulating and optimizing a robust language model, the expandable strategy to ensemble the well-trained single models. Besides, we conducted a meticulous evaluation of these techniques, and a comprehensive comparison of the experiments indicated the effectiveness of these techniques in video captioning.