Pretrained Image-Text Models are Secretly Video Captioners
Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi · 2025
Chunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush Vosoughi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). 2025.