Semi-Automatic Aligning of Swedish Forensic Phonetic Phone Speech in Praat using Viterbi Recognition and HMM
Jonas Lindh · 2007
Automatic alignment of text and sound is of great help and saves a lot of time labelling speech databases, either for research or for developing speech technology tools such as automatic speech recognition or text to speech systems. It is also a very useful tool in forensic speaker identification as one often receives a tapped recording together with an orthographic transcription. The orthographic transcription can be used together with the sound file to provide information of where in the recording significant events occur to a greater or lesser extent. Even if the aligning sometimes is not perfect, it replaces some of the time consuming manual labelling. To perform automatic aligning, common speech recognition techniques are applied at various levels. In this case, a framework for doing automatic aligning, called EasyAlign, was developed for the free software Praat (Goldman, 2007). Praat is distributed as an open source software under a GPL license. On top of the source code a built-in scripting language can execute commands, make calculations and communicate with other programs in different manners (Boersma & Weenink, 2007). To be able to implement a new language for automatic aligning within the framework there is a need for a grapheme to phone converter and a trained Hidden Markov Model that can be used by the viterbi recognition program HVite from the HTK toolkit (Young et al., 2006). Automatic Aligning of Speech Aligning recorded speech automatically is a technique that borrows heavily from automatic speech recognition (ASR). Successful attempts have been made using Hidden