Applying Specialised Corpora to Language Testing : Development of the Exam Corpus (Vocabulary and Grammar Sections)
Hiroko Usami · Institutional Repositories DataBase (IRDB) · 2021
A corpus is an electronically collected written or spoken text that represents a particular language for the purpose of linguistic analysis (e.g., Baker, Hardie, & McEnery, 2006;McEnery, Xiao, & Tono, 2006).As a methodology, corpora have been applied to language teaching topics such as computer-assisted language learning (CALL), datadriven learning (DDL), and compiling teaching materials (Chambers, 2010; Cheng, 2010, Chapter 23; Walsh, 2010, Chapter 24), among others.It has been suggested that corpora can play an important role in language testing (Cushing, 2017;Park, 2014).Perhaps the most basic application for corpora in language testing is the use of specialised corpora containing academic and field-specific English terms that have been compiled for the main purpose of developing wordlists and test items.In contrast, there have been few corpora that contain past test items to inform, validate, and assess the English used in examinations, and to analyse aspects of grammar or vocabulary frequently tested in examinations.Therefore, the aim of this paper was to introduce an original, specialised corpus created by the author for language testing, the Exam Corpus.Currently, this corpus contains 1,191,850 words used in multiple-choice vocabulary and grammar questions in different worldwide English proficiency examinations.This paper describes the design, data, metadata collection, annotation, and application of the Exam Corpus after reviewing previous studies on applying corpora to language testing, especially about utilising specialised corpora in this area.