Can AI Provide Useful Holistic Essay Scoring?
Tamara Powell Tate, Jacob Steiss, Drew H. Bailey, Steve Graham, Daniel Ritchie, Waverly Tseng, Youngsun Moon, Mark Warschauer · 2023
Researchers have sought for decades to automate holistic essay scoring. Over the years, these programs have improved significantly. However, accuracy requires significant amounts of training on human-scored texts—reducing the expediency and usefulness of such programs for routine uses by teachers across the nation on non-standardized prompts. This study analyzes the output of multiple versions of ChatGPT scoring of secondary student essays from three extant corpora and compares it to quality human ratings. We find that the current iteration of ChatGPT scoring is not statistically significantly different from human scoring, but exact agreement with humans is still difficult. Consistency and agreement within one point, however, is achievable and may be sufficient for low-stakes, formative assessment purposes.