Learning second language speech perception in natural settings
Emily R. Felker · Radboud Repository (Radboud University) · 2021
Chapter 1: IntroductionLearning how to listen in a second language is a crucial component of second language acquisition and is vital for successful communication.While speech perception in a native language (L1) is usually an effortless process, second language (L2) speech perception poses many difficulties for the non-native listener, the most notable of which is word recognition (Cutler, 2012).One of the challenges is hearing the difference between two speech sounds that are contrastive in the L2 but not in the L1 (Best, 1994;Best & Tyler, 2007).For instance, native Dutch listeners may confuse the English words "pan" (/pæn/) and "pen" (/pɛn/) since they perceptually assimilate both /æ/ and /ɛ/ into the same phonemic category in their L1.Another challenge for L2 listeners is adapting to regional or foreign accents.For example, a Dutch traveler in Australia may confuse the words "pen" (/pɛn/) and "pin" (/pɪn/) because they are not familiar with the /ɛ/-to-/ɪ/ vowel shift in certain Australian English dialects.An important question in second language acquisition is how L2 listeners can overcome these word recognition challenges to improve their perception of L2 speech.Traditionally, L2 speech processing and perceptual learning have been studied using non-interactive, computer-based training programs that provide intensive exposure to de-contextualized, highly controlled stimuli (e.g., Bradlow et al., 1999; Iverson & Evans, 2009; Lively et al., 1994; Logan et al., 1991; Sakai & Moorman, 2018).These experimental paradigms have yielded valuable insights about how various aspects of the phonetic input can contribute to perceptual learning.However, they do not address the more natural learning situations that L2 listeners encounter in everyday life, such as learning in conversational interaction.Reconciling control over phonetic input with ecological validity is an important methodological goal for L2 speech learning research: this would allow us to test whether certain L2 speech learning mechanisms, proposed on the basis of research using artificial listening tasks, transfer to real-life communicative settings.This dissertation aims to make methodological and theoretical contributions to the study of improving L2 speech perception in natural contexts.As an introduction, Section 1.1 will first motivate the need for improving the ecological validity of L2 speech perception research, focusing on the importance of more naturalistic speech stimuli and learning contexts.Section 1.2 will introduce three perceptual learning mechanisms, ranging from relatively implicit to relatively explicit, that are investigated throughout the dissertation: lexical guidance, interactional corrective feedback, and phonetic instruction.Finally, Section 1.3 will present a chapter-by-chapter outline specifying the research questions and methodologies addressed in this dissertation. Expanding the scope of research to more natural language processingIn recent years, researchers in psycholinguistics and related fields have called for studying language processing of more natural types of speech in more ecologically valid contexts (e.g., Tanenhaus & Brown-Schmidt, 2008;Tucker & Ernestus, 2016; Willems, 2017).One way to achieve this is to study the processing of continuous speech, rather than isolated words, and to use stimuli from more casual speech registers, such as conversational speech.Conversational speech differs from formal speech not only in syntactic structure and lexical content but also in numerous phonological and phonetic properties (e.g., Tucker & Ernestus, 2016).In particular, reduced pronunciation variants, in which phonemes are weakly articulated or altogether missing, are a hallmark of casual speech (e.g., Ernestus & Warner, 2011;Johnson, 2004) and have been shown to influence the processes involved in speech perception (e.g., Brouwer et al., 2012;Janse & Ernestus, 2011;Kemps et al., 2004).These reduced pronunciation variants often cause speech comprehension problems for L2 listeners (e.g., Brand & Ernestus, 2018;Ernestus et al., 2017).Experiments that employ recordings of continuous speech extracted from real conversations would represent a step forward in the direction of studying the type of speech that L2 listeners are likely to encounter in everyday life.To investigate the perception of continuous speech, dictation tasks that require listeners to transcribe the words they hear have long been used in speech intelligibility research (Kent, 1992) and in L2 teaching (Field, 2003;Morris, 1983;Oller & Streiff, 1975;Siegel & Siegel, 2015;Stansfield, 1985).However, dictation tasks are typically only scored for accuracy at the word level (Buck, 2001; Kent, 1992;Savignon, 1982), thereby excluding information about what phonemes in particular posed problems for the listener.Only recently have more fine-grained scoring measures, such as those based on phonetic feature similarity between a listener's transcription and a target phrase, come into use in phonetics research (e.g., Podlubny et al., 2018).Chapter 2 demonstrates how Chapter 2: Evaluating dictation task measures for the study of speech perceptionAbstract: This paper shows that the dictation task, a well-known testing instrument in language education, has untapped potential as a research tool for studying speech perception.We describe how transcriptions can be scored on measures of lexical, orthographic, phonological, and semantic similarity to target phrases to provide comprehensive information about accuracy at different processing levels.The former three measures are automatically extractable, increasing objectivity, and the middle two are gradient, providing finer-grained information than traditionally used.We evaluate the measures in an English dictation task featuring phonetically reduced continuous speech.Whereas the lexical and orthographic measures emphasize listeners' word identification difficulties, the phonological measure demonstrates that listeners can often still recover phonological features, and the semantic measure captures their ability to get the gist of the utterances.Correlational analyses and a discussion of practical and theoretical considerations show that combining multiple measures improves the dictation task's utility as a research tool.