Research on Wheat Seed Classification Based on Machine Learning Algorithms and Data Analysis Visualization

Kaili Sun, Wei Bai, Jiexin Feng, Zhe Yang, Yanyan Li · Advances in Computer Signals and Systems · 2025

This study addresses the problem of wheat seed classification by employing three machine learning algorithms—Random Forest (RF), Naïve Bayes (NB), and Support Vector Machine (SVM)—on the Wheat Seeds Dataset from the UCI database. Through comprehensive data preprocessing, feature analysis, and model construction, the impact of different feature combinations on classification accuracy was systematically investigated. The dataset, comprising 210 samples with seven attributes (e.g., area, perimeter, and kernel groove length), was standardized and split into training and testing sets to ensure robust evaluation. The experimental results demonstrate that RF and SVM significantly outperform NB in classification performance, with SVM achieving the highest accuracy of 97.61% when combining area or width with kernel groove length. Notably, the combination of perimeter and kernel groove length yielded the highest accuracy (96.67%) in RF, while compactness and asymmetry coefficient consistently performed poorly across all algorithms, with accuracy as low as 60.71% in SVM. These findings highlight the critical role of feature selection in classification tasks, with kernel groove length emerging as a key determinant. This research not only provides an effective technical reference for wheat variety classification but also underscores the practical value of machine learning in agricultural applications, offering insights for optimizing efficiency and reducing costs in food security initiatives.

Read the paper · More papers on PaperTik