Pretrained Image-Text Models are Secretly Video Captioners

Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi · 2025

Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). 2025.

Read the paper · More papers on PaperTik