A Comparative Analysis of Domain Adaptation Techniques for Recognition of Accented Speech

Gyurgy Szaszaak, Piero Pierucci · 2019

Addressing the domain mismatch problem has been a long term interest within cognitive infocommunication. In particular, several techniques to compensate for dialectal variations of the same base language in Automatic Speech Recognition (ASR) have been proposed. Conservative retraining, transfer learning, multi-task training, matrix factorization, i-vector based techniques as well as adversarial and teacher-student training, have been proposed for the specific purpose of ASR deep neural acoustic models domain adaptation. Comparing these techniques is often complicated as different experiments are carried out on diverse datasets and within various frameworks. It is also worthwhile analyzing possible combination of such techniques within complex systems. The objective of this work is to systematically compare and analyse a number of domain adaptation techniques for ASR using the same framework, the open-source Kaldi toolkit, in order to allow for a fair comparison on adapting US English acoustic models for the Indian accent. Our results indicate that, when properly hyper-parametrized and carefully regularized, the easiest approaches, requiring less complexity and reduced computational power, can perform equally well as the more complex ones.

Read the paper · More papers on PaperTik