Prompt Tuning Empowering Downstream Tasks in Multimodal Federated Learning
Yichen Bao · 2024
Federated learning (FL) has made it possible to train global models using decentralized data in a privacy-preserving manner through the aggregation of model updates. Current FL methods have been expanded to incorporate multi-modal data for training a global model across multiple data modalities. However, global model performs poorly on downstream tasks. Recently, Prompt engineering has emerged as an indispensable technique for extending the capabilities of large language models (LLMs) and vision-language models (VLMs). This approach leverages task-specific instructions, known as prompts, to enhance model efficacy without modifying the core model parameters. Since prompt engineering is only applied to pre-trained language models and has not been widely used in other domains, in this paper, we propose a framework to apply prompt engineering from language models to multi-modal federated learning. We design prompts for representative downstream tasks including image classification task, text classification task, and multimodal retrieval task. Then, we freeze the backbone of the global model and do a few-shot sample study on prompts. Afterwards, we empirically analyze our framework via extensive experiments, and show its superiority in performance.