A new approach to the analysis and annotation of speech and prosody based on computerized cross-linguistic corpora.
Dolores Ramírez Verdugo · Procesamiento del lenguaje natural · 2003
In the present paper, corpus linguistics becomes a valuable methodological tool for cross-linguistic research on speech and prosody. The inherent complexity of speech analysis and prosodic annotation increases when the object of study is a longitudinal computerized corpus of native and nonnative varieties of English. The lack of generally accepted prosodic transcription systems adds further difficulty to the task. The fact that the most popular transcription models such as INTSINT or ToBI are intended for Standard varieties of a language led us propose a multidimensional level of annotation which has proved to be effective in identifying the non-native speakers’ main prosodic characteristics and their implications in the discourse. I present a new approach to the acoustic, and discourse analysis of two computerized corpora of non-native and English native language speakers (460 hours of spoken language, 3.32.400 words). These parallel corpora belong to an on-going longitudinal research project: the UAM Corpus of Spoken English as a Second Language, funded by Autonomous Community of Madrid (CAM, 06/0027/2001). I survey and annotate the prosodic patterns used by both native and non-native language speakers aiming at describing the extent to which the intonation systems used by non-native speakers may affect the information structure and the discourse meaning of their messages. I propose new levels of annotation as Figures 1 and 2 illustrate. NSS // 1 But/ last No/ vember/ wasn’t/ cold// //1Was it//