Analysis of BERT-Based Federated Learning in Text Classification: A Study on Data Distribution and Algorithm Selection

Zaixi Jia · 2025

Federated learning has emerged as a popular distributed machine learning framework designed to preserve privacy. However, in non-IID (non-Independent and Identically Distributed) training data environments that more closely resemble real-world scenarios, the performance of models trained through federated learning can be adversely affected. This paper explores the application of a BERT-based federated learning framework in Chinese news classification tasks, with a focus on analyzing the impact of different data distributions on model performance and the selection strategies of federated learning algorithms. We compare the performance of FedAvg and FedProx algorithms in both IID and non-IID (label distribution skewed) data environments, and investigate the influence of the proximal term hyperparameter$\mu$in FedProx on model convergence and performance. The research results indicate that: models trained under IID data distribution generally outperform those under non-IID data distribution; in non-IID data distribution scenarios, the FedProx algorithm demonstrates better stability and accuracy compared to FedAvg; appropriate adjustment of the proximal term hyperparameter$\mu$in FedProx can further optimize model performance in non-IID environments. This study provides empirical references for the application of federated learning in Chinese natural language processing tasks and offers guidance for addressing data heterogeneity issues in practical applications.

Read the paper · More papers on PaperTik