Predicting Success in CS1 - An Open Access Data Project
Keith Quille, Keith E. Nolan · Proceedings of the 53rd ACM Technical Symposium on Computer Science Education V. 2 · 2022
PreSS# is an online Machine Learning prediction model that aims to identify students at risk of failing or dropping out in an introductory programming course (typically called CS1). PreSS# has been developed over the past 16 years, where the model is capable of predicting at-risk students with an accuracy of 71%. There is, however, a need to re-validate the model using a larger international multi-jurisdictional multi-university data set, as up until now the data sets have been predominantly from a single jurisdiction. The goal of this study is to not only re-validate the model using a multi-jurisdictional data set, but, in line with a 2015 ITiCSE working group report's Grand Challenges, to openly publish the data set itself. This work timely to the CSEd community as other researchers can use this data to further their research, re-validate PreSS# and will be able to then contribute, by submitting their local PreSS# data sets to this global online repository.