Distributed Random Forests for NIDS with Edge and Global Model Aggregation
Mohammed Maruf Hossen, Mohamed Rafi, Tazrian Alam · 2025
The growing sophistication of cyber threats necessitates scalable, privacy-preserving solutions for network intrusion detection. This study proposes Federated Forest, a decentralized framework leveraging federated learning (FL) to collaboratively train a global Random Forest model across distributed edge nodes without sharing raw data. The architecture enables local nodes to preprocess data (handling missing/infinite values via imputation, normalizing features, and applying SMOTE for class balancing) and train localized Random Forest classifiers, enhanced by Isolation Forest for outlier detection. Hyperparameters (e.g., tree depth, estimators, split criteria) are optimized per node using GridSearchCV, tailored to local computational constraints. Model parameters (feature importance, tree structures) are aggregated at a central server via weighted averaging, with weights assigned based on node performance, data quality, and resource capacity. Evaluated on the CICIDS2017 dataset, the global model achieves 98.5% accuracy, 98.42% precision, and 98.1% F1-score, with local nodes (Random Forest Models 1–4) demonstrating 99.49–99.71% precision. The framework maintains stable performance across heterogeneous nodes while preserving data privacy through encrypted parameter transmission and GDPR-compliant aggregation. Key innovations include (1) a federated aggregation mechanism for tree-based models, (2) node-specific hyperparameter tuning, and (3) preprocessing pipelines addressing class imbalance and outliers in decentralized settings. The results validate the viability of federated Random Forests as a scalable, privacy-aware solution for collaborative intrusion detection in distributed networks.