Automated evaluation of essays and short answers
Jill C. Burstein, Claudia Leacock, Richard D. Swartz · Loughborough University Institutional Repository (Loughborough University) · 2001
Essay questions designed to measure writing ability, along with open-ended questions requiring short answers, are highly-valued components of effective assessment programs, but the expense and logistics of scoring them reliably often present a barrier to their use.Extensive research and development efforts at Educational Testing Service (ETS) over the past several years (see http://www.ets.org/research/erater.html) in natural language processing have produced two applications with the potential to dramatically reduce the difficulties associated with scoring these types of assessments.The first of these, e-rater™, is a software application designed to produce holistic scores for essays based on the features of effective writing that faculty readers typically use: organization, sentence structure, and content.The e-rater software is "trained" with sets of essays scored by faculty readers so that it can accurately "predict" the holistic score a reader would give to an essay.ETS implemented e-rater as part of the operational scoring process for the Graduate Management Admissions Test (GMAT) in 1999.Since then, over 750,000 GMAT essays have been scored, with e-rater and reader agreement rates consistently above 97%.The e-rater scoring capability is now available for use by institutions via the Internet through the Criterion Online Writing Evaluation service at http://www.etstechnologies.com/criterion.The service is being used for both instruction and assessment by middle schools, high schools, and colleges in the U.S. ETS Technologies is also conducting research that explores the feasibility of automated scoring of short-answer content-based responses, such as those based on questions that appear in a textbook's chapter review section.If successful, this research has the potential to evolve into an automated scoring application that would be appropriate for evaluating short-answer constructed responses in online instruction and assessment applications in virtually all disciplines.