Data mining at the intersection of psychology and linguistics
R. Harald Baayen · Max Planck Digital Library · 2005
Large data resources play an increasingly important role in both linguistics and psycholinguistics.The first data resources used by both psychologists and linguists alike were word frequency lists such as Thorndike and Lorge (1944) and Kučera and Francis (1967).Although the Brown corpus on which the frequency counts of Kučera and Francis were based was very large for its time, comprising some one million word forms carefully sampled from different registers of English, many common words did not appear in the frequency lists, while others appeared with counterintuitive frequencies of use.Gernsbacher (1984) addressed this issue, claiming that subjective frequency estimates would be superior to objective frequency counts.Corpus-based frequency counts would be inherently unreliable due to regression towards the mean.In another corpus, higher frequency words would be less frequent, and lower frequency words would be more frequent.These considerations have led many psychologists to turn away from research directly addressing frequency effects in lexical processing.This distrust in psychology of corpus-based frequency data mirrors the rejection of corpora as a valid source of information about grammar in generative linguistics.Fortunately, more and larger corpora were developed, driven in part by the needs of commercial lexicography, in part by the research interests of corpus linguistics, and in part by the growing needs for reliable data in computational linguistics and linguistic engineering.These develop-5 70 BAAYEN ments made the creation possible of the CELEX lexical database, an initiative of the psycholinguist Levelt, which is widely used in both the (computational) linguistic and psycholinguistic research communities.For English, this resource provides the frequencies in the Cobuild corpus at the time that this corpus comprised some 18 million words.The British National Corpus (BNC) currently is the largest available tagged corpus of British English, with 100 million words, of which 10 million transcribed spoken English.Thus, linguistics now has at its disposal large data resources, although much remains to be done with respect to annotation and the sampling of everyday spoken language.The largest unstructured source of examples of language in use is, nowadays, the World Wide Web, which combines the advantage of quantity with the disadvantages of the absence of linguistic annotation and the restriction to written language.The lexical resources developed specifically within psychology are relatively new, scarce, and small compared to linguistic corpora.Perhaps the most important large data resources are WordNet (Miller, 1990;Fellbaum, 1998), the Florida association norms (Nelson, McEvoy, & Schreiber, 1998), and the databases of visual lexical decision latencies, word naming latencies, and subjective frequency ratings of Balota, Cortese, and Pilotti (1999) and Spieler and Balota (1998).These resources provide psycholinguistics with a wealth of data on the behavioral properties for thousands of words.Although here too much remains to be done, especially from a morphological point of view, these behavioral data resources are a tremendous step forwards compared to the small numbers of items typically studied in factorial psycholinguistic experiments.The aim of this chapter is to show that, when combined, the linguistic and psychological resources become a particularly rich gold mine for the study of the lexicon and lexical processing.I will illustrate the new methodological possibilities for data mining by examining the databases compiled by Balota and colleagues, in combination with CELEX, the BNC, and WordNet.For 1424 monomorphemic and monosyllabic nouns, and 832 monomorphemic and monosyllabic verbs, we study the predictive potential of a range of variables for three behavioral measures: visual lexical decision latencies and word naming latencies in ms, and subjective familiarity ratings on a 7-point scale.In what follows, I will show that mining these combined resources yields several new insights.Section 1 examines the correlational structure of the predictors, and sheds new light on the nature of word frequency.Section 2 shows that subjective frequency ratings are an independent variable in their own right, just as response latencies in, for instance, DATA MINING IN PSYCHOLINGUISTICS