Automated Scoring for NAEP Short-Form Constructed Responses in Reading
Mark D. Shermis · 2024
This study was commissioned by the National Assessment of Educational Progress (NAEP) to determine the feasibility of using machine prediction scoring algorithms to evaluate reading items for fourth- and eighth-grade reading. The study consisted of two parts and took place as a set of open competitions. The first part (prompt-specific) asked competitors to score a test set of 20 items drawn from the 2017 NAEP administration. Twelve competitors were provided item materials, a randomly selected training and validation set, and asked to make predictions for a test set where only the text of the response was provided. In the second competition (generic), five competitors were asked to score two reading items where no training or validation set was provided. For the test set, competitors had only access to the text of the response but could draw upon information from other sources. In the first competition, the human rater performance was κ ω = 0.91; for the competition winners, the scoring performance was κ ω = 0.89, 0.88, and 0.87. In the generic competition, the winner had a κ ω = 0.53.