Automated Measurement of Vowel Formants in the Buckeye Corpus - eScholarship
Yao Yao, Sam Tilsen, Ronald Sprouse, Keith Johnson · 2010
UC Berkeley Phonology Lab Annual Report (2010) Automated Measurement of Vowel Formants in the Buckeye Corpus Yao Yao 1 , Sam Tilsen 2 , Ronald L. Sprouse 1 , and Keith Johnson 1 University of California, Berkeley University of Southern California Abstract: In recent years, corpus phonetics has become a rapidly expanding field. However, the lack of appropriate tools for automatic acoustic analysis hinders further development of the field. In this paper, we present a methodological study on the automatic extraction of vowel formants using both robust linear predictive coding (RLPC; Lee, 1988) and dynamic formant tracking (Talkin, 1987). Acoustic data were taken from the Buckeye corpus of English conversations. We varied two aspects of the analysis - preemphasis and LPC order - to optimize formant tracking results by speaker and vowel. We also show, based on the optimal results, the distribution of ten English vowels in the F1/F2 space in conversational speech. Keywords: speech corpus, automatic acoustic analysis, vowel formants, Robust LPC 1. Introduction 1.1. Background With the development of various speech corpora in recent years (e.g. Switchboard corpus of telephone speech, TIMIT speech database, Buckeye corpus of conversational speech), a new trend has emerged in the field of phonetic research, which involves large-scale quantitative analysis of acoustic corpus data. Compared with experimental methods, this new line of research, which is often termed “corpus phonetics”, features the use of larger and more realistic phonetic datasets as well as more sophisticated data analysis. Over the past couple decades, the new methodology has produced a fast growing body of literature on various topics including phonetic variation (Byrd 1994; Keating et al, 1994; Raymond et al, 2006; Bell et al., 2003, 2009, etc), speech tempo (Fosler-Lussier and Morgan, 1999; Yuan et al., 2006; Jacewicz et al., 2009, etc), disfluency in natural speech (Shriberg, 2001, etc) and so on. A classic paper by Keating et al. (1994) presented two case studies using the TIMIT speech database. One of their studies was based on the non-audio part of the corpus (i.e. speaker information, temporal and segmental transcriptions) and investigated durational variation and vowel alternation in the pronunciation of the word the. The other study analyzed audio data from the corpus and tested the effect of following vowels on the articulation of velar stops. However, this balanced use of audio and non-audio data has not been pursued in later studies, as non-audio transcriptions have been far more frequently consulted than audio data. As a result, most corpus analyses have concentrated on segmental duration and alternation, but little has been learned about more fine-grained phonetic detail, such as VOT and vowel formants.