Automated Scoring of Speaking Tasks in the Test of English‐for‐Teaching ( TEFT ™)
Klaus Zechner, Lei Chen, Larry Davis, Keelan Evanini, Chongmin Lee, Chee Wee Leong, Xinhao Wang, Su‐Youn Yoon · ETS Research Report Series · 2015
This research report presents a summary of research and development efforts devoted to creating scoring models for automatically scoring spoken item responses of a pilot administration of the Test of English‐for‐Teaching (TEFT™) within the ELTeach™ framework. The test consists of items for all four language modalities: reading, listening, writing, and speaking. This report only addresses the speaking items, which elicit responses ranging from highly predictable to semipredictable speech from nonnative English teachers or teacher candidates. We describe the components of the system for automated scoring, comprising an automatic speech recognition (ASR) system, a set of filtering models to flag nonscorable responses, linguistic measures relating to the various construct subdimensions, and multiple linear regression scoring models for each item type. Our system is set up to simulate a hybrid system whereby responses flagged as potentially nonscorable by any component of the filtering model are routed to a human rater, and all other responses are scored automatically by our system.