Defining New Words in Corpus Data: Productivity of English Suffixes in the British National Corpus
Eiji Nishimoto · eScholarship (California Digital Library) · 2004
The present study introduces a method of identifying potentially new words in a large corpus of texts, and assesses the morphological productivity of 12 English suffixes, based on some 78 million words of the written component (books and periodicals) of the British National Corpus (BNC).The method compares two corpus segments (created by randomly sampling at the level of documents within the BNC), and defines new words as those that are not shared across segments (segments being interpreted as randomly sampled speaker groups).The approach taken differs from others in the literature in that new words are identified irrespective of how many times a given word is used by the same speaker (author).A productivity ranking of the 12 English suffixes is obtained, and the results are shown to be intuitively satisfying and stable over different sample sizes.With a psycholinguistic interpretation of the data, implications for the nature of intuitions about productivity are considered.