A Gaussian Mixture Model and Mutual Information Guided Dataset Partitioning Algorithm for Convolutional Neural Network Training

Lida Shahmiri, Patrick Wong, Laurence S. Dooley · 2025

Traditionally, dataset partitioning for training, testing, and validation purposes in Convolutional Neural Networks (CNNs) has used either random selection or cross-validation techniques. However, in complex datasets such as the VNPlant-200 medicinal plant collection of images in their natural habitat, high intra-class and low inter-class variations between captured images make it challenging to reliably achieve accurate classification. Mutual Information (MI) has recently been applied as a similarity measure to guide the CNN training set selection process by synthesising more representative images. While the MI-Guided Training (MIGT) algorithm improved accuracy performance, misclassification rates for certain species were still high due to the recurring influence of extraneous background clutter, blurring and illumination effects in many of the images. This paper introduces a novel Gaussian Mixture Model-Mutual Information Guided Training (GMM-MIGT) dataset partitioning algorithm that uniquely incorporates a combination of Expectation-Maximisation (EM), Gaussian Mixture Model (GMM) clustering, and Principal Component Analysis (PCA) as a two-stage preprocessing pipeline before MI-based partitioning. GMM-MIGT lowers colour complexity, removes less useful dimensions, and enhances similarity estimation, ensuring more representative partitions. Experimental results validate the GMM-MIGT algorithm's consistently superior performance over existing partitioning methods, achieving classification accuracies between 83% and 99%, allied with significantly lower misclassification rates, especially in complex datasets with high visual variability.

Read the paper · More papers on PaperTik