A Study on Forced Alignment Error Patterns in Kaldi
A K Punnoose · 2022 8th International Conference on Signal Processing and Communication (ICSC) · 2022
Forced alignment is one of the fundamental pieces in speech recognition, and we generally overlook the errors in the forced alignment process. It is extremely time consuming to look for phoneme level or word level alignment errors, and most importantly the alignment error detection is subjective to the listener. In this paper, we seek to identify error patterns in forced alignment, particularly the alignment skewness. We define an alignment skewness scoring function to capture the overall skewness of the alignment. Four different hypotheses, the first one on the overall skewness of the forced aligner, the second one on the difference in forced alignment between clean speech and noisy speech, the third one on male vs female recordings, a fourth one on the alignment difference between the word preceding silence vs random word, are formulated. One sample t-test for the first hypothesis and two sample t-test for the next three hypotheses are conducted respectively on a Kaldi trained forced alignment on Voxforge dataset. We report a left skew on the overall forced alignment output compared to the manually labelled ground truth and, no statistically significant difference in alignment skewness between clean speech vs noisy speech, male vs female alignment, and the word preceding silence vs random word alignment.