SweVoc - A Swedish vocabulary resource for CALL
Katarina Heimann Mühlenbock, Sofie Johansson Kokkinakis · 2012
The core in language teaching and learning is vocabulary, and access to a delimited set of words for basic communication is central for most CALL applications. Vocabulary characteristics also play a fundamental role for matching texts to specific readers. For English, the task of grading texts into different levels of difficulty has long been facilitated by the existence of lists serving as guides for vocabulary selection. For Swedish, the situation is with a few exceptions less fortunate, in that no base vocabulary organized according to aspects of usage has existed. The Swedish base vocabulary – SweVoc – is an attempt to remediate this. It is a comprehensive resource, aimed at differentiating vocabulary items into categories of usage and frequency. As we are of the opinion that no corpus of written text can do fully justice of general language use, we have utilized materials from a second language as reference for delimiting the category of core words. Another belief is that the task of defining a base vocabulary can not be fully automatic, and that a considerable amount of manual, traditional lexicographic work has to be invested. Hence, the present approach is not an innovative, but a methodological approach to list generation for a specific purpose, much like LSP. We anticipate SweVoc to be integrated in CALL applications for vocabulary assessment, language teaching and students’ practice. 1 Background Vocabulary knowledge plays a central role in a person’s ability to communicate, as well as reading and understanding written text. It is therefore a central issue in many readability assessment approaches. Prominent researchers within readability and language assessment, such as (Thorndike, 1921; Vogel and Washburne, 1928; Patty and Painter, 1931; Thorndike and Lorge, 1944; Dale and Chall, 1948; Spache, 1953), and more recently (Nation, 1990; Nation, 2001), all included specific lists as a criterion to measure text difficulty for English. In quantitative associative studies of readability, some scheme for measuring the vocabulary difficulty is set up, compared to a predefined criterion, and expressed by a coefficient of correlation. In this way, the lists may be constructed in order to mirror vocabulary difficulty corresponding to school grade levels. Thorndike’s (1921) list of 10,000 words, later on revised into a list of 30,000 words (Thorndike and Lorge, 1944) and Spache’s revised list (Spache, 1974) of 1,040 entries, were mainly constructed by judgment and common sense. West published in 1953 the General Service List – a list of 2,000 words selected to represent the most frequent words in an English corpus. Vocabulary is also an important issue when producing language-supportive aids for persons with deficient communication capability. Insufficient vocabulary knowledge implies a decrease in expressive power of an utterance or written text, and the receptive language skills are also heavily dependent upon the individual vocabulary range. In order to obtain maximum benefit from language supportive tools, the resources provided as lists ought to be chosen with care in order to conform to individual and situational needs. Also in generating LSP (language for specific purposes) and particular domain vocabulary lists, a list of general base vocabulary is needed in order to exclude the most common and general words. In the following we are making a distinction between base vocabulary and core vocabulary. A language teaching situation might involve a more extensive base vocabulary, while assistive technology applications such as symbol boards for communication would benefit from a restricted core vocabulary, expandible with complementary vocabulary items from different domains. The present approach is an attempt to combine both models, i.e. it is a Swedish core vocabulary list, supplied with words belonging to a broader base vocabulary. Defining a core vocabulary is a task associated with several methodological challenges. Lee (2001) has enumerated some of them. First of all, the concept of core vocabulary has to be settled. Several working definitions exist, out of which the most contested point seems to be whether the list is based on, and intended for, applications within written or spoken language, or both. If one decides to adopt the view that a core vocabulary is by definition that which is central to the language as a whole, it rules out for instance approaches based on frequency countings of words in written language. Furthermore, it should be untarnished from any stains of genre, style, register or lect association. In addition to the theoretically founded issues, also problems of more practical nature arise. Although a major part of verbal communication is said to take place with the use of 1,500 2,000 words (West, 1953), this figure must be considered in the light of language-specific properties, of the type of communication, and above all, as a function of the concept. Counting lexemes, lemmas, baseform orthographic words or multiwords render different figures. For English, the notion of plays a central role when defining list for educational purposes. Lee (2001), citing Schmitt (2000) maintained that people in the field seem to agree that the word family is the most meaningful unit to work with and pedagogically most useful. The concept was put forth by Bauer and Nation (1993), from a reader’s perspective defined to comprise a base and all its derived and inflected forms that can be understood by a learner without having to learn each form separately. If all the lemmas belonging to a specific are considered as one member of the list, Hirsh and Nation (1992) found that a vocabulary size of at least 5,000 entries were needed in order to read unsimplified fiction texts. The same study also showed that graded readers beginning at a level of 2,600 families would be of great benefit in language teaching. An attempt to construct a levelled base vocabulary for another language than English was made by De Mauro (1980) when he published a list of 7,400 Italian words, categorized into three different groups according to use. The only attempt in this direction for Swedish was made by Forsbom (2006), who derived a base vocabulary pool from a corpus of 1 million words – the Stockholm-Umea Corpus (SUC) (Kallgren, 1992). This was achieved by ranking base forms according to adjusted frequency over the entire corpus, and then adopting a subsequent filtering technique that sorted out entries which did not occur in more than three out of nine genres in the corpus. The result was a Swedish base vocabulary pool (henceforward referred to as SBVP), with a total amount of ≈ 8,200 base forms, mirroring the use of written Swedish in the early nineties. SBVP alone neither be considered to reflect modern language use, nor to be enough informative to independently serve as a source of words pertaining to a restricted core vocabulary, since it is based solely on written language. As already mentioned, the base forms in SBVP are ranked according to adjusted frequency (AF: see equation 1), i.e. relative frequency weighted with dispersion over the 9 categories (genres) in SUC. It implies that the vocabulary are those words that are not genre dependent, given the subdivisions of a small-size text corpus. Furthermore, it lacks information at the lexeme level, which reduces its feasability for purposes demanding a semantic disambiguation between words. A base form like the Swedish noun gang has for instance four lexeme representations, belonging to different base vocabulary categories. The first refers to ’time’ and is considered to be a core vocabulary item, while the sense ’path’ is not. The second Issues regarding a distinction between lemma and lexeme concepts are discussed in Gardner (2007). Another flaw in SBVP is the absence of internal levelling, which would be required in order to serve as a list of core vocabulary words. In the present approach, it was hence enriched with labels indicating levels of general use from three additional sources; (1) a translated base vocabulary, (2) a list of words from modern vocabulary, and (3) a dictionary of words denoting domestic life activities and participation in community activities. The final product is SweVoc, a base vocabulary list, consisting of ≈ 8,500 words, mainly lemma forms, divided into five different categories.