Improvement of Elderly Speech Recognition Using Gammatone Filterbank Adaptation

Kazumasa Yamamoto, Akinori Ishiki, Seiichi Nakagawa · 2021 IEEE 10th Global Conference on Consumer Electronics (GCCE) · 2021

Recently, the accuracy of speech recognition has been remarkably improved by the introduction of deep learning. However, the accuracy of speech recognition for old-old/oldest-old people is still insufficient and a spoken dialogue system for them has not been put into practical use. In this study, GtFDNN-HMM, which uses a gammatone filterbank for the feature extraction layer in the DNN, is used as an acoustic model, and speaker/elderly adaptation is performed to improve the accuracy of speech recognition for the old-old/oldest-old people. Only the parameters of the gammatone filterbank are applied for adaptation of GtFDNN, so speaker/elderly adaptation with a small amount of speech is possible, and it is suitable for speech recognition for old-old/oldest-old people, where it is difficult to collect a large amount of speech. From the result of experiments, the speech recognition performance was improved by using the elderly speech adaptation model.

Read the paper · More papers on PaperTik