Feature Scaling Optimization for Durian Characteristic Prediction using Machine Learning
Thanaphon Phukseng, Jakkaphun Nanuam, Pattharaporn Thongnim · 2025
Durian is one of the most economically valuable fruits in Southeast Asia particularly in Thailand where accurate yield prediction is essential for market planning and resource management. This study investigates the effect of feature scaling techniques on the performance of machine learning models in predicting durian yield using cultivation area and geographic information. A Random Forest model is employed due to its robustness and ability to handle both numerical and categorical data. Three feature scaling methods including Min-Max Scaling, Z-score Standardization and Robust Scaler are evaluated in combination with two encoding strategies which are Label Encoding and One-Hot Encoding. Experiments are conducted under two data splitting approaches which are Random Split and Time-Based Split simulating real world scenarios. The model is evaluated using Mean Absolute Error, Root Mean Squared Error and R-squared. Results show that Z-score Standardization consistently yields the highest predictive accuracy when combined with Label Encoding and time-based data splitting. Label Encoding also slightly outperforms One-Hot Encoding in both accuracy and computational efficiency. This research highlights the importance of proper preprocessing in agricultural prediction tasks and recommends a preprocessing pipeline of Label Encoding, Z-score Standardization and time-based evaluation for optimal durian yield prediction using Random Forest.