EMERGING TECHNOLOGIES : Tools and Trends in Corpora Use for Teaching and Learning

Bob Godwin-Jones · Language learning & technology · 2001

INTRODUCTION Language corpora have long been exploited for language instruction. Vocabulary lists for learners, for example, have been generated from corpora, and word counts derived from corpus analysis have helped in defining goals for vocabulary acquisition. Dictionary and textbook creators have used corpora extensively. In recent years, the move to the use of authentic language materials in language pedagogy has enhanced the role collections of spoken or written language can play in language learning. Corpora are, after all, huge storehouses of real language use. The interest in languages for special purposes further favors the use of corpora, as a means to identify the specific language components to be taught. Technology enhancements have made corpora more widely available, as well as provided more powerful tools for their use. In particular, the Internet is playing a steadily growing role in the dissemination of corpora and corpus-based teaching materials. Corpora are no longer the exclusive domain of lexicographers and computational linguists. ACCESS TO CORPORA Corpora are of interest today to professionals in a wide variety of fields, from ethnologists to telecommunication conglomerates. Creating a language corpus is a major undertaking, both time-consuming and expensive. This is all the more the case for collections which include multiple languages and/or audio/video recordings. Given the cost and the growing interest, it makes little sense for corpora not to be made widely accessible. In fact, there have been a large number of corpora in many different languages which have become available over the Internet in the last few years. Good starting points for finding them are Michael Bohman's Corpus Linguistics page, the Linguistic Exploration page (at the LDC - Linguistic Data Consortium) or the Tractor page (the Telri Research Archive of Computational Tools and Resources). These pages in many cases link to direct corpus access, including a number of parallel corpora of particular interest in translation studies and language learning for specific purposes. There are as well a substantial number of text collections of literary works in a variety of languages. Some include comprehension aids and annotations for use in language learning. As the number of language archives grows, locating the specific resources needed for a project will become more problematic. One can only go so far with lists of Web links (even when annotated) or traditional Web searching. There is a recently launched international project, the Open Language Archives Community (OLAC), to build an infrastructure linking language archives of all types together. OLAC builds on the Open Archives Initiative and on the Dublin Core Metadata Initiative. The Dublin Core project began in 1995 to develop conventions for resource searching on the Web. OLAC uses the core 15 elements of the Dublin Core and extends them through the use of qualifiers to fit the needs of the language community. The use of a controlled vocabulary of descriptors should allow more efficient searching of archives. The consistent use of meta-data in language resources is likely to become of growing importance in the language community. There has not been a standard way to include information about a resource, such as the participants in an interview included in a corpus (i.e., age, nationality, first language, education, etc.). Such information is typically included in a header which is either part of the resource file itself or stored separately. Because of the different ways such meta-information has been stored, there has been a proliferation of tools and approaches for the user to access that information. It would be very helpful for both researchers and users to have a common approach to resource description, not only for corpora, but for all language resources such as text collections, lexicons, grammar tutorials, multimedia files, Web lessons, and so forth. …

Read the paper · More papers on PaperTik