Federated Learning for Network Traffic Classification: Impact of Non-IID Distribution on Model Performance
Yi-Hsien Chiang, Lin-Huang Chang, Tsung-Han Lee · 2023
Network traffic classification is vital for network security. However, the non-IID nature of participant data poses challenges for federated learning in this context. This study examines the impact of non-IID on model performance in federated learning for network traffic classification. We analyze three non-IID data distribution scenarios: distribution-based label imbalance, quantity-based label imbalance, and quantity skew. We evaluate the effect of these distributions on federated learning performance using training models with varying degrees of non-IID data. Our experiments evaluate model accuracy and convergence efficiency. We also compare the influence of different non-IID levels on model performance. Results show that federated learning effectively handles non-IID challenges and performs well across all scenarios. However, increased data heterogeneity or label imbalance reduces model accuracy and convergence efficiency. Models perform relatively poorer with non-uniform label distributions. In conclusion, federated learning is a viable solution for non-IID in network traffic classification, emphasizing the need to consider the impact of data heterogeneity. These insights can guide the application of federated learning in network traffic classification and related fields.