Performance Evaluation of Indonesian Language Forced Alignment Using Montreal Forced Aligner

Griffani Megiyanto Rahmatullah, Shanq-Jang Ruan · 2023

Manual segmentation is crucial for aligning words or phonemes between speech signals and their transcription. It is used in many applications, such as speech recognition, speaker diarization, and speech synthesis. However, because manual segmentation can be laborious and time-consuming, forced alignment algorithms are used to automate the process by aligning the speech audio to its corresponding text transcription at the phoneme or word level. To overcome this problem, one of the solutions is a forced alignment tool called Montreal Force Aligner. This tool has two requirements: a language acoustic model and a corresponding language dictionary. The challenges are determining a suitable recipe for training the model and achieving a good alignment with minimum error. In this study, training and evaluation are conducted for the Indonesian language using the Montreal Forced Aligner. A conversational corpus from a range of five to 30 speakers with various duration between one hour and two hours is used as an input to train the acoustic model. Then, the model is employed to align the conversational Indonesian corpus using the ASR-INDOCSC test set. Our findings include a recommendation training strategy for developing an Indonesian language acoustic model using the Montreal Forced Aligner tool, which measures the number of speakers and hours required.

Read the paper · More papers on PaperTik