Vision language distillation by clustering bitrajectory matching

Jiaming Zhou, Shangjiaqi Hao, Qinghao Zhang · 2024

Dataset distillation is often used to create compact datasets that can be used to achieve similar training performance, making it a good choice for addressing the challenges of data storage cost and training cost. However, existing distillation method are generally time-intensive and computationally expensive, especially when applied to vision-language tasks. To address this challenge, we propose the Clustering BiTrajectory Matching method, which accelerates existing distillation techniques by 8 times through two innovative strategies: a clustering-based sample selection and a biTrajectory optimization approach. The Clustering BiTrajectory Matching method can achieve good accuracy in a multi-modal setting while requiring lower computation resources and emphasizing efficiency in pre-training. We evaluate the proposed method on the Flickr8k dataset. We show that our method is able to achieve better efficiency (less iteration to achieve target accuracy), while outperforming other coreset selection methods.

Read the paper · More papers on PaperTik