The Impact of Training Data on Automated Short Answer Scoring Performance

Michael Heilman, Nitin Madnani · 2015

Automatic evaluation of written responses to content-focused assessment items (automated short answer scoring) is a challenging educational application of natural language processing.It is often addressed using supervised machine learning by estimating models to predict human scores from detailed linguistic features such as word n-grams.However, training data (i.e., human-scored responses) can be difficult to acquire.In this paper, we conduct experiments using scored responses to 44 prompts from 5 diverse datasets in order to better understand how training set size and other factors relate to system performance.We believe this will help future researchers and practitioners working on short answer scoring to answer practically important questions such as, "How much training data do I need?"

Read the paper · More papers on PaperTik