Model Selection Through Cross-Validation for Supervised Learning Tasks with Manifold Data

Derek Brown · Journal of Purdue Undergraduate Research · 2024

Research Snapshots 123of a model.The algorithm starts by partitioning the data into K-folds.The model is first trained on all but one fold, with its accuracy being evaluated on the left-out fold.This is repeated for each fold, with the overall model performance being evaluated by taking the average of each of these accuracies.Notably, each accuracy is correlated with the others since the training folds overlap.This correlation makes establishing large sample asymptotic theorems, such as the central limit theorem, for K-fold cross-validation difficult.Recent work has started to establish central limit theorems for K-fold cross-validation for real-valued random variables.My research expands these theorems to include the case in which random variables represent angles.Similar to how a unit circle can be defined with a single angle, multiple simultaneously recorded angles lie on a higher dimensional space called a torus, which looks like a doughnut in three-dimensional space.Such spaces are of relevance in biology.For example, in biomechanics, data extracted from locomotion experiments involve joint angles, and in biochemistry, backbone dihedral angles of amino acid residues determine protein structures.Understanding K-fold cross-validation for angular valued variables will provide valuable tools to scientists wanting to use accurate machine learning algorithms in many different areas involving data with complex structures. Model Selection Through Cross-Validation for Supervised Learning Tasks with Manifold Data Student researcher:Derek Brown, Senior K-fold cross-validation is a popular method used in machine learning for estimating the prediction accuracy

Read the paper · More papers on PaperTik