Classifying Programming Ability: A Machine Learning Approach Using Debugging Metrics
Basma S. Alqadi · 2024
The high failure rate in introductory programming courses (CS1) remains a persistent challenge, prompting research into predictive models that can identify struggling students early. This study investigates the potential of using debugging performance metrics as predictors of programming ability in novice programmers. We examine the statistical association between debugging performance -measured by correctness levels and debugging time- and overall programming ability, as well as the effectiveness of machine learning models in classifying students as high or low achievers based on these metrics. Using data collected from students enrolled in a CSI course, we employ Student's t-test to analyze this association and find that correctness levels are significantly higher among high achievers. We then train and evaluate three machine learning models-Random Forest, k-Nearest Neighbors (k-NN), and Binary Logistic Regression-using 1,000 iterations of out-of-sample bootstrap validation. Our results demonstrate that all models show promising predictive performance, with Random Forest achieving the highest accuracy (0.86), recall (0.81), precision (0.92), and F1 score (0.85).