Collection and detailed transcription of a speech database for development of language learning technologies
Harry Bratt, Leonardo Neumeyer, Elizabeth E. Shriberg, Horacio Franco · 1998
We describe the methodologies for collecting and annotating a Latin-American Spanish speech database. The database includes recordings by native and nonnative speakers. The nonnative recordings are annotated with ratings of pronunciation quality and detailed phonetic transcriptions. We use the annotated database to investigate rater reliability, the effect of each phone on overall perceived nonnativeness, and the frequency of specific pronunciation errors. 1. INTRODUCTION In this paper we describe the methodologies for collecting and annotating a Latin-American Spanish speech database. The database includes recordings by native and nonnative speakers. A panel of listeners rated the pronunciation quality of the nonnative data, and a group of expert phoneticians phonetically transcribed a subset of the nonnative data. The database was intended for use in the development of hidden-Markov model (HMM) based speech technologies for language learning [4], including robust speech recognition...