Federated Learning for Medical Applications - A Study on Performance and Bias with Logistic Regression on Small Datasets
Kirthika Ashokkumar, Sakina Rahman, Palak Agarwal, Mahima Agumbe Suresh · 2024
Important concerns when using medical data for Machine Learning (ML) is patient privacy and bias. Federated Learning (FL), the training of a centralized model by using parameters from decentralized models, is alternatively used to protect patient privacy. Medical data can often be structured and sparse, where deep learning is not applicable. In this work, we applied Federated Learning with Logistic Regression on medical data through multiple experiments with data distribution among clients. Three simulated cluster sampling methods conducted to compare model accuracy with different data distributions and/or sample sizes. Our observations were as follows: (1) the Federated Learning model performs, on average, better than the average of individual clients, and (2) variability increases as the sample size and the number of clients increases, and accuracy stays moderately consistent. In addition, we designed a Federated Learning system that securely transfers model parameters between server and clients. We also present an approach to study bias in these models. Our results show that Federated Learning is a promising approach for medical applications.