Using prosody in automatic segmentation of speech

Ossama Essa · 1998

Automatic speech segmentation is an essential tool for building large corpora for training speech recognition systems.Manual segmentation of speech is both time consuming and an error-prone task.Several automatic segment,ation systems have been proposed based on the acoustical features of the speech [5] [9].In rhythmic speech, prosodic features become an essential factor in designing an accurate automatic segmentation system.This work presents a novel t,echnique for automatic segmentation of speech in which both prosodic and acoustical features of the speech are examined to achieve a higher accuracy of segmentation.The system was tested on Koranic Arabic, a highly rhythmic language.This paper shows that incorporating the prosodic features in the design resulted in better segmentation accuracy for rhythmic speech.Low center unrounded vowel High back rounded vowel High front unrounded vowel Voiced lottal stop Voiced rlabial stoo E Unvoiced dental s&p Unvoiced inter-dental fricative Voiced dental fricative Unvoiced pharyngeal fricative Unvoiced velar fricative Voiced dental stop Voiced inter-dental fricative Voiced dental trill Voiced dental fricative Unvoiced dental fricative Unvoiced palatal fricative Unvoiced velarized dental fricative Voiced velarized palatal stop Unvoiced velarized palatal stop Voiced velarized interdental fricative Voiced pharyngeal fricative Voiced uvular fricative Unvoiced labiodental fricative Unvoiced uvular stop Unvoiced velar stop Voiced dental sonorant Voiced bilabial nasal Voiced dental nasal Unvoiced

Read the paper · More papers on PaperTik