Automated Essay Scoring and The Repair of Electronics
Dan Preston, Danny Goodman · 2012
The Hewlett Foundation sponsored the Automated Student Assessment Prize on kaggle.com, challenging teams to produce essay evaluation models that best approximate human graders. Contestants predicted the scores of standardized-testing essays from grades 7-10. Teams were provided with 8 sets of labeled training data. Each set corresponds to a different essay prompt, grading rubric, and range of possible scores. In general, grading rubrics cite content, fluidity, spelling, grammar, and vocabulary as major considerations. Teams submitted predicted grades for an unlabeled test set, and the results were then ranked by a complicated scoring metric called mean quadratic weighted kappa. The contest concluded on April 30, 2012. Automated essay grading is a difficult domain because computers struggle with several of the tasks an expert human performs during grading, such as implicitly correcting syntax mistakes and evaluating the logical structure of an argument. The problem is compounded by the poor agreement of individual human graders. Computers have traditionally approached the problem by applying machine learning techniques to a mix of simple features extracted from the essays, which do not capture a human’s intuition about what ought to cause a good grade. Past work has been done with student essays written in good faith; that is, the students expected human evaluators. If future students are aware that computers are grading their work, some will attempt to write for the simple heuristic features alone, without constructing a coherent essay. The evaluation system must guard against this tactic. In this paper, we present our current essay evaluation model for the Kaggle contest, and our plans for future work.