Imbalanced Datasets and Crop Yield Prediction: Application of Preprocessing Techniques for Regression Tasks in Agriculture

Mariaelisa Polsinelli, Morteza Mesbah, Zhiming Qi, Matt Ramsay · 2024

Abstract. Machine learning (ML) is increasingly used in the agricultural sector to predict crop yields. The ability to produce short-term seasonal forecasting or future projections is highly important for creating climate change adaptation strategies. The escalation of extreme weather events such as drought due to climate change is a significant threat to crop yields. The increase in weather extremes presents a challenge as historical data may lack instances of low yield years, leading to unbalanced datasets biased towards ‘normal‘ or ‘well performing‘ years. Three preprocessing techniques; random undersampling (US), random oversampling (OS), and the Synthetic Oversampling Technique for Regression (SMOTER) were applied to the training datasets for predicting the yields of nine industrial farm field-years growing Russet Burbank using random forest (RF). US was the winning technique for five of the nine field-years, likely due to reducing overfitting, and when assessing all nine field-years overall (Baseline: 16.7%, US: 14.1%, OS: 17.5%, SMOTER: 16.1% RRMSE). SMOTER provided the largest improvement in performance for the drought field-year (Baseline: 36.8%, US: 31.1%, OS:31.1%, SMOTER: 26.2% RRMSE). The application of SMOTER to synthetically increase low yield data points and high growing degree days (GDD)/low precipitation values in the training dataset reduced the RRMSE to satisfactory (20% ≤ RRMSE < 30%) from poor (≥30%). The results indicate US may help reduce overfitting and SMOTER can be a viable option for improving the yield prediction of drought and extreme weather years, potentially in combination with other performance techniques.

Read the paper · More papers on PaperTik