Predicting student performance in introductory computer programming courses

William E. J. Doane · 2008

For decades, computer science education researchers have sought to improve computing education by refining curricula, instructional methods, and choice of first language. Central to the task of improving computer science education is the identification of students in need of assistance, ideally as early in their academic career as possible. Until recently, no known assessment instrument offered a good predictor of student performance in introductory computer programming courses. Such an instrument, should it be created, would allow educators to identify students who would be likely to have difficulty learning to program. It would also allow instructors to design instruction intended to support those students, and to allocate instructional resources more appropriately. In 2006, researchers in the United Kingdom identified an assessment instrument that shows promise as a predictor of students' final grades in introductory computer programming courses. In that same year, researchers in Massachusetts found that the commercially available logic puzzle MasterMind® also showed promise as a predictor of in-class programming test scores. What connects these techniques? What makes them more successful than past assessment instruments designed to test programming potential? In this study, novice programmers at the undergraduate and high school levels completed a modified version of the paper-based assessment instrument designed by U.K. researchers. Students also were asked to complete web-based tasks based on MasterMind® and Sudoku. The purpose of this study was to collect information about novice computer programming students and to use that information effectively to predict their final numeric course grade in introductory computer programming courses. In order effectively to extract all of the most relevant information from the initial collected data, both model-based and algorithmic prediction methods were used in the predictive analyses. Regression trees were used, in addition to model-based multiple regression methods, to derive both generalizable and interpretable predictive results, given the available data sets.

Read the paper · More papers on PaperTik