Enhancing Soil Fertility Prediction Through Federated Learning on IoT-Generated Datasets with a Feature Selection Perspective
Murali Krishna Senapaty, Abhishek Ray, Neelamadhab Padhy · 2024
Introduction: Fertile soil has a balanced pH and nutrient profile (potassium, phosphorus, and nitrogen), water retention capability, and organic substances. Fertile soil allows for better plant growth, leading to better production. The soil fertility requirements vary from crop to crop. So, it is essential to identify the soil fertility level according to the crop type. Objective: The objective of this paper is to develop a robust model that is capable of predicting the soil fertility. The model is integrated with IoT-generated data and federated learning-based feature selection techniques to improve the accuracy of the dataset. Materials/Methods: Different feature selection techniques were applied to the dataset. Then, we applied machine learning algorithms such as logistic regression, decision tree, and naïve Bayes, as well as their combinations to analyze and improve the performance. The federated learning approach was implemented to train the local models using the individual partitioned datasets. Each local model of the client shared the cryptic output weight and bias without sharing the raw data. There was a centralized model at the server end that collected these weights and biases, preserving data privacy. These collected data were aggregated and applied to find the least square error (LSE). Then, a gradient descent curve (GDC) was applied to identify the optimized weight and bias, which were fed back again to improve the accuracy of the predictions. Result: From our experimental observations, we analyzed the performance metrics of different ML classifiers, and it was revealed that the ensemble of logistic regression and decision tree had a better performance than the other models. One of our client models generates weight and bias with a precision of 87%, an accuracy of 87%, a recall of 87%, and an F1-score of 86%. Further, we collected two of our client system model outcomes from a server model and applied the LSE to identify the optimal W and B. In future work, we wll improve the performance of our model with a recursive approach by verifying the W and B at the client model in a feedback process.