Effective Utilization of Pre-Trained Models in Ad-Hoc Video Search

Kazuya Ueki, Yuma SUZUKI, Takayuki Hori, Hiroki TAKUSHIMA, Hideaki OKAMOTO, Hayato Tanoue · Journal of the Japan Society for Precision Engineering · 2023

In recent years, image and text retrieval in the zero-shot learning setting has increasingly become more accurate by utilizing pre-trained image and text embedding models. By adapting these models to the task of video retrieval, we were able to achieve the world's highest retrieval accuracy in the 2022 ad-hoc video search task of TRECVID, a video retrieval and evaluation benchmark conducted by the National Institute of Standards and Technology. This paper reports on an investigation and experiments to determine the effectiveness of a number of trained image and text embedding models currently available to the public for video retrieval tasks, using a large set of over 1.4 million test videos used in the TRECVID benchmark.

Read the paper · More papers on PaperTik