Phonetic labeling and acoustic correlates for building Japanese speech data base

Yushinori Sagisaka, Shigeru Katagiri, Kazuya Takeda · The Journal of the Acoustical Society of America · 1987

Fine description of a large amount of speech signals is indispensable to acquire effective acoustic-phonetic rules in speech synthesis and recognition. To build a finely labeled speech data base, manual labeling is carried out using digital sound spectrograms and additional acoustic parameters that reflect power and spectral characteristics. Through labeler training, labeling items are modified several times to decrease the deviation of segment boundaries and to hasten labeling speed. As a result, not only the usual phonemic categories, but also finer phonetic events (e.g., closure, burst, and aspiration for plosive consonants) are labeled. Moreover, multiple segment boundaries (e.g., boundaries between vowels and following fricative consonants) and inseparable portions (e.g., aspiration followed by a devocalized vowel) are specially marked. To ensure labeling quality, reliability tests and error analysis are carried out. Labeling criteria using these results will be used for a large scale data base construction.

Read the paper · More papers on PaperTik