Fully Decentralized Collaborative Learning for Visual Question Answering in Distributed Scenarios
Yudai Ueda, Hideya Ochiai · 2025
With the rapid advancement of Large Language Models (LLMs), foundation models are increasingly used to finetune specific tasks. In some cases, they are personalized to specific situations, such as particular input domains. Simultaneously, Federated Learning (FL) has gained attention as an effective approach to preserve training data privacy. In particular, fullydecentralized FL-a form of collaborative learning framework leveraging multiple devices with peer-to-peer communicationprovides a promising architecture for model personalization. In this paper, we propose WAFL-VQA, an initial implementation of fully decentralized FL for Visual Question Answering. We validate its capability to effectively personalize models based on each device's training data. In particular, we focused on the indirect acquisition of information about images unavailable to individual devices and the relationship between the number of model parameters and performance. Our experiments demonstrate that validation accuracy improved by up to 30% except for one device whose performance remained at the same level. This highlights that incorporating information from other devices can enhance model performance. Furthermore, experiments with low-rank adaptation (LoRA) suggested the potential to reduce the number of trainable model parameters.