Assessing the processing consequences of segment reduction in Dutch with naive discriminative learning

R. Harald Baayen · Lingue e linguaggio · 2010

This study addresses the comprehension of reduced words, taking as point of departure two lexical decision experiments reported in Ernestus (2009). Ernestus discusses the consequences of segment reduction in auditory comprehension in terms of exemplars for reduced forms and generalization processes reconstructing the unreduced form. A different approach is explored in the present study, using a computational model based on discriminative learning to explain the pattern of results in the experimental data. This new modeling approach, which provides the researcher with detailed information into what distributional properties of the language input may drive the observed effects, suggests that the unusual biphones in reduced words are the key to understanding why reduced words can be learned and understood. In a recent study, Ernestus (2009) investigated the processing of Dutch past participles (e.g., ge-vraag-d, ‘asked’), which are formed by simultaneously prefixing ge([x@]) and suffixing -d or -t. In Dutch, the schwa in unstressed prefixes such as geis often omitted. Ernestus presented speakers of Dutch with the past participles of phonotactically legal new monomorphemic verbs. In a familiarization phase, these past participles were combined with pictures in order to ensure that the nonce verbs received an interpretation. A week later, participants were asked to complete an auditory lexical decision experiment in which these past participles were presented. One question addressed in this study was whether reduced forms acquire their own exemplar representations (c.f., e.g., Johnson, 2004) facilitating subsequent processing. A second question was whether the priviliged status of unreduced forms (Ernestus, Baayen, & Schreuder, 2002; Gaskell, 2003) would affect the speed of comprehension. In Experiment 1, participants were exposed either to the reduced or to the full form of the new participles during familiarization, and later performed a lexical decision on the reduced forms. These two conditions will henceforth be referred to as R+R (reduced for training and reduced for testing) and as U+R (unreduced for training, reduced for testing). I am indebted to Mirjam Ernestus for making the words she used in her experiments available to me, and to both Mirjam Ernestus and Vito Pirelli for their insightful comments on an earlier version of this paper. SEGMENT REDUCTION AND NAIVE DISCRIMINATIVE LEARNING 2 Response latencies were significantly shorter for the R+R condition (1271 ms) compared to the U+R condition (1319 ms). In theories framed in terms of form representations, the results of Experiment 1 can be interpreted as evidence for the coming into existence of form representations for reduced words after familiarization. When subsequently encountered in the lexical decision task presenting reduced forms as targets, the newly-formed representations of reduced forms would then provide a better match to the acoustic input, and therefore would give rise to shorter latencies compared to the unreduced words, which would have form representations that do not fully match the input. Experiment 2 used the same familiarization procedure, but presented unreduced forms (instead of reduced forms) as targets in the lexical decision experiment. The mean latencies for these R+U (reduced training, unreduced testing) and U+U (unreduced training, unreduced testing) conditions were very similar (1328 ms and 1330 ms respectively). Ernestus argued that when listeners encounter a new reduced form, the unreduced form is also reconstructed from the reduced form. Hence, when the unreduced form is presented in lexical decision, it matches a representation irrespective of whether an unreduced or a reduced form was presented during familiarization. As a result, average response latencies for the U+R and U+U conditions are indistinguishable. This interpretation raises several questions. First, if two representations come into existence upon encountering a novel reduced form, some theories predict competition between these representations (e.g., Luce & Pisoni, 1998). In the R+R and R+U conditions, in which representations for both the reduced and unreduced forms are supposedly available after familiarization, a processing delay would be expected. No such delay is present in the experimental data, however. In fact, the R+R condition leads to shorter instead of longer response latencies than the U+R condition. Second, the reconstruction of the unreduced form from the reduced form is supposedly driven by a generalization giving canonical status to the unreduced form. In hybrid models combining rules and exemplars (see, e.g., Atallah, Frank, & O’Reilly, 2004; Goldinger, 2007), large-scale generalizations are supposed to be early processes, whereas episodic, exemplar-driven generalization is supposed to be a late process, manifesting itself primarily in elongated processing times. However, when the two experiments of Ernestus (2009) are considered jointly, it is the R+R condition, the condition in which slow exemplardriven processing should be most prominently involved, for which we observe the shortest latencies, instead of the longest latencies. Third, positing representations for both unreduced and reduced forms comes with the risk of a proliferation of representations, given the high degrees of variability characterizing speech. Finally, whereas modeling the processing of acoustic information as an early process and the processing of speaker-specific episodic information as a late process (Goldinger, 2007) may have its advantages. Information about, for instance, a speaker’s voice plays no role in phonological or phonetic generalizations about the characteristics discriminating one word from the other words in the lexicon. Yet episodic information about a speaker’s voice may co-determine lexical processing at, if Goldinger is correct, later stages in the comprehension process. The present experimental data, however, concern specific and general information that is qualitatively quite similar. Instead of episodic information about a speaker’s voice in conjunction with segmental information, the data compare the presence SEGMENT REDUCTION AND NAIVE DISCRIMINATIVE LEARNING 3 versus absence of a segment in a lexical representation. Qualitatively similar differences permeate the lexicon, for any pair of words that differ in the presence versus the absence of a segment (e.g., hand, hands and hand, had. Although it is possible that the exemplardriven and generalization-based processes posited to explain the present experimental data are taking place at different sites and at different points in time, this possibility becomes less attractive when both processes concern the same kind of segmental information. In this study, therefore, a very different approach to understanding these experimental data is pursued, using a computational model first proposed in Baayen, Milin, Filipovic Durdjevic, Hendrix, and Marelli (2010). This model makes use of naive discriminative learning based on the Rescorla-Wagner equations (Wagner & Rescorla, 1972; Danks, 2003), which are well-established in psychology as a mathematical description of learning. In what follows, a brief introduction to this computational model is presented first. Next, the model is illustrated by pitting its predictions for Dutch inflected verb forms against the by-item mean latencies available in the Dutch Lexicon Project, henceforth dlp (Keuleers, 2010). We then zoom in on the Dutch past participle, and the relative importance of prefix and verb stem for lexical access as measured by the lexical decision task. Finally, simulations are presented clarifying that the pattern of results observed by (Ernestus, 2009) follows straightforwardly from discriminative learning. Naive discriminative learning The model developed by Baayen et al. (2010) is a two-layer network with symbolic representations for a word’s form and a word’s semantics. A word’s form is coded by means of its unigrams and bigrams (or uniphones and biphones). Its meaning is represented in terms of the semantic units associated with its constituents. Thus, the word bookcases is linked with the meanings book, case, plural. Each input unit is linked to each meaning unit, and a weight is associated with each link. When a word is read (or heard), the activations of its unigrams and bigrams are set to 1, and those of all other unigrams and bigrams to 0. Activation is then propagated through the connections. The activation ai of a meaning i is defined as the sum of the weights on its active incoming links: ai = ∑

Read the paper · More papers on PaperTik