jTRACE: A Reimplementation and Extension of the TRACE Model of Speech Perception and Spoken Word Recognition

Harlan D. Harris, James S. Magnuson, Ted J. Strauss · eScholarship (California Digital Library) · 2005

jTRACE: A Reimplementation and Extension of the TRACE Model of Speech Perception and Spoken Word Recognition Ted J. Strauss ([email protected]) James S. Magnuson* ([email protected]) Harlan D. Harris ([email protected]) Department of Psychology, University of Connecticut, 406 Babbidge Road, Unit 1020 Storrs, CT 06269 USA Abstract * recognition in real-time. TRACE takes the main ideas of Cohort theory and implements them within the Parallel Distributed Processing framework (McClelland & Rumelhart, 1981), thus allowing the model to be implemented as a computer program, and generate detailed, falsifiable predictions about human behavior. TRACE successfully models many aspects of real-time spoken language processing, including categorical perception of phonemes, word segmentation, lexical effects on phoneme processing, lexical competition (McClelland & Elman, 1986), and even fine-grained time course data from eye tracking (e.g., Allopenna, Magnuson, & Tanenhaus, 1998; Dahan, Magnuson, & Tanenhaus, 2001; Dahan, Magnuson, Tanenhaus, & Hogan, 2001), among many other phenomena (see Protopapas, 1999, for a review of TRACE and its place in the literature, and Frauenfelder & Peeters, 1998, for analyses of various TRACE parameters). This paper describes jTRACE, a freely-available, cross- platform reimplementation of the TRACE model of speech perception and spoken word recognition in Java. The goal of the reimplementation is to facilitate the use of simulations by researchers who may not have the skills necessary to use the original C implementation of TRACE. In this paper, we report a large scale validation project, in which we have replicated a number of important previous simulations, and then we describe several new features in jTRACE designed to help researchers conduct original TRACE research as well as to replicate earlier findings. These include visualization tools, powerful scripting, built- in data analysis and graphing, stochasticity, and save/load functions that facilitate archiving and sharing simulations. Overview TRACE (McClelland and Elman, 1986) is arguably the best psychological model of speech processing to date, as it is able to simulate the deepest and broadest range of empirical phenomena. However, it is not widely used. One obstacle is that the original implementation in the C programming language is opaque to the average psychologist (or even the average programmer). We present jTRACE, a user-friendly, cross-platform and free software tool that reimplements the TRACE model in the Java programming language. Researchers at different technical levels have different modeling needs. jTRACE accommodates most of these needs, hiding details from the beginner, giving powerful scripting tools to the advanced user, and providing a basis for easy extensibility by programmers. The introduction gives a primer on TRACE and the motivations for creating this tool. The second section reviews some old and new simulations that we have replicated with jTRACE. The third section describes the principal functions that make this an effective and versatile tool. Readers are encouraged to download jTRACE from: How TRACE Works The TRACE model is a connectionist network with an input layer and three processing layers: feature, phoneme and word. There are bottom-up excitatory connections between input-feature, feature-phoneme and phoneme-word layers. There are lateral inhibitory connections between units within the feature, phoneme and word layers. There are top-down excitatory (feedback) connections between word- phoneme and phoneme-feature layers (although phoneme- feature feedback is typically set to 0.0; McClelland & Elman, 1986). An external stimulus is given to the input layer and on each processing cycle, and activation passes along the weighted connections, changing the activation values of units in the processing layers. The input to TRACE is a pseudo-spectral representation (McClelland & Elman, 1986). The input takes the form of a 63-dimensional vector at each time increment. Each time increment in TRACE is intended to approximate about 10 milliseconds of real time, and the 63- dimensional vector describes the activation of 7 acoustic features, each comprised of 9 continua. Feature units have excitatory connections to phonemes. TRACE uses a 14 phoneme subset of English phonemes plus a “silence” phoneme: /p/, /b/, /t/, /d/, /k/, /g/, /s/, /S/, /r/, /l/, /a/, /i/, /u/, /^/, and the silence phoneme, /-/. Phoneme units have excitatory connections to word units. Simulations can be run with no items in the lexicon, just a few, or hundreds, depending on simulation needs (simulations are faster with smaller lexicons). The original, standard lexicon (slex) has 212 items, though fairly large lexicons have http://maglab.psy.uconn.edu/jtrace.html Introduction In their seminal paper, McClelland and Elman (1986) describe TRACE and successful modeling of human behavior on a wide array of speech processing tasks. Prior to TRACE, the Cohort model (Marslen-Wilson & Tyler, 1980) laid the groundwork for modeling spoken-word Also at Haskins Labs, New Haven, CT.

Read the paper · More papers on PaperTik