What and How Much Evidence Do We Need? Critical Considerations in Validating an Automated Scoring System
Xiaoming Xi · 2008
Building on Clauser, Kane and Swanson (2002), this paper illustrates how an argumentbased approach can be applied to the validation of the TOEFL ® iBT Speaking test which uses an automated scoring system called SpeechRater v.1.0. The paper outlines assumptions pertaining to the links between each stage in the score interpretation and decision making process. Finally, evidence needed to reject potential rebuttals against the inferences is described. By outlining the inferences underlying score interpretation, the paper shows the connections among various aspects of validity evidence and offers insights into practical issues arising in a validation process such as prioritization of different types of evidence.