Automatic alignment of phonetic segments
Kåre Sjölander · 2009
Speech data transcribed at the phoneme level is important for basic speech technology applications. This paper describes some experiments with an automatic method for aligning a given sequence of phonemes with the corresponding spoken utterance. It is shown that methods borrowed from the field of automatic speech recognition can successfully be adapted to this problem. Results are reported for experiments that have been carried out on speech data collected in the Waxholm and SWEDIA 2000 projects. In order to confirm the performance level and generality of the proposed method, corresponding experiments have been conducted on the American English TIMIT corpus. For the test utterances from the Waxholm data the best system positions 85.1% of all boundary locations within 20 ms of the manually segmented reference boundaries. For the test utterances from the SWEDIA material, the corresponding figure is 70.9%.