Effects of cognitive load in speech production: an experimental corpus combining prosody, syntax and eye-tracking data

George Christodoulides · Digital Access to Libraries · 2015

We present an experiment designed to constitute a corpus of speech produced under varying levels of cognitive load, by manipulating task complexity and demands on working memory. Cognitive load (CL) is a multidimensional phenomenon defined as the amount of mental demand imposed by a particular task on the performer (Paas, 2003), or as the perceived effort invested by her during the execution of the task. CL is the result of a task placing demands on cognitive systems with limited capacity, such as working memory (Baddeley 2007). Since cognitive overload negatively affects performance and may induce performance errors, the ability to estimate CL in real-time can be very useful, especially in high-stress environments. Pupillometry (Just et al. 2003) is a non-intrusive psycho-physiological method to quantify cognitive load, based on the observation that cognitive activity is correlated with pupil dilation (cf. Chen & Epps 2012). It has been used to study CL induced by language perception (e.g. Demberg et al. 2013, Engonopoulos et al. 2013; Kun et al. 2013). A group of 11 university students (4 M, 7 F, French mother-tongue) participated in the experiment, which included four phases. Phase 1 consisted of a Stroop test under time pressure. In Phase 2, participants read aloud a short text (self-paced reading, one sentence at a time), and a subsequent multiple-choice comprehension question; they selected one out of four possible answers, justifying their choice by recalling information from the short text just read (10 text-question-response triplets were recorded). In Phase 3, self-paced reading of a narrative text (a fictitious newspaper article often used in French phonology studies) was followed by 3 comprehension questions (summarise the text, explain its main idea, and state two additional points). In Phase 4, self-paced reading of an argumentative text on economic policy was followed by the same 3 comprehension questions, with the addition of distractors (strings of numbers that the participants heard through headphones at random intervals, and had to recall after providing their answer to each comprehension question). Finally, a short interview with the experimenter was conducted, during which participants were encouraged to speak freely. The objective of this experimental design was to collect speech produced under increasing levels of load to the working memory, including stretches of utterances long enough to analyse both prosodic features and syntactical / discourse segmentation strategies. The corpus total length is approximately 12 hours, and it is comparable with the one for English presented in (Yin et al., 2007, 2008). Speech was recorded with a high-precision head-mounted microphone, while we simultaneously collected eye tracking data (gaze and pupil size) using the Pupil (Kassner & Patera 2012) portable eye tracker. This time-synchronised multi-track recording is transcribed and aligned to the phone level, allowing us to extract a series of segmental and prosodic features using a cascade of semi-automatic tools. The features studied include temporal characteristics such as the distribution of silent pauses and speech rate; the distribution and prevalence of disfluencies; prominent syllables, including their patterning and density; pitch range and intonation patterns; prosodic phrasing and its relationship with syntactical phrasing (resulting from an automated syntactical analysis). The findings are compatible with previous research comparing speech production under normal and high cognitive load conditions, including research on simultaneous conference interpreting (Christodoulides 2013). It has been shown that cognitive load, in the form of increased demands placed on working memory subsystems, affects pause duration and distribution, articulation rate and disfluencies (Berthold & Jameson 1999; Müller et al 2001; Jameson et al. 2009), as well as phonetic and prosodic features (Yin et al. 2008; Tet Fei Yap 2012, Petrone et al. 2011). The analyses on correlating the prosodic features with syntax, and observations from eye-tracking data are continuing and will be presented in detail in the conference.

Read the paper · More papers on PaperTik