Robust Speech-Age Estimation Using Local Maximum Mean Discrepancy Under Mismatched Recording Conditions

Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama, Yusuke Ijima · 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) · 2021

A recently proposed time-delay neural network (TDNN)-based age estimation system has yielded state-of-the-art performance in speech-age estimation tasks. However, the performance of this TDNN-based system can seriously degrade when the recording conditions of each utterance are different in the training and testing phases. To tackle this problem, we examine the efficiencies of a series of unsupervised domain adaptation (UDA) methods to obtain the model invariance against the difference of these conditions. In particular, we propose using local maximum mean discrepancy (LMMD) with soft-target labels to consider an ordinal relationship between age labels. In most UDA methods, the model is trained to obtain domain invariant representations by minimizing the statistical difference of the distributions between labeled source and unlabeled target data without considering their age class labels. In contrast, our LMMD-based approach locally minimizes the differences in their distributions on each age class while considering adjacent age classes using soft-target labels. We conducted speech-age estimation experiments on in-house datasets under mismatched conditions including different background noise, reverberation, and microphones. The experimental comparison demonstrated that the LMMD-based method contributed to efficiently reducing the effect of mismatches of input data, yielding significant improvements over other UDA methods, such as MMD and reverse gradients.

Read the paper · More papers on PaperTik