Automated Identification and Validation of the Optimal Number of Knowledge Profiles in Student Response Data
Bradley P. Din, Tanya Nazaretsky, Yael Feldman-Maggor, Giora Alexandron · 2023
It is well--known that the provision of personalized instruction can enhance student learning. AI--based education tools can be used to incorporate blended learning in the science classroom, and have been shown to enhance teachers' ability to prescribe this personalisation. In order to reveal student knowledge profiles from their response data, we must utilise classical educational data mining techniques, with cluster analysis being one of the key methods for doing so. However, while clustering algorithms typically require the number of clusters as a hyperparameter, there is no clear method for choosing the optimal number.Motivated by a practical instance of this foundational problem -- deciding on the number of student clusters for a group-based personalization tool -- this paper discusses several variations of the gap statistic to identify the optimal number of clusters in student response data. We start with a simulation study where the ground truth is known to evaluate the quality of the identified methods, and then assess their behaviour on real student data. Based on the results, we suggest a stability--based approach to validate our predictions, and identify an empirical threshold for the number of observations for a prediction to be stable.We find that if a dataset has some cluster structure, very small subsamples also showed cluster structure -- large datasets were not required to identify if any cluster structure exists, but only to discern the number of clusters accurately. Finally, we discuss how the flexibility of the method enables teachers to choose between a different number of clusters according to their class environment or teaching goals.