Whispered Speech to Normal Speech Conversion Using Bidirectional LSTMs with Meta-network
Wei-Wei Yu, Hailun Lian, Jian Zhou, Huabin Wang, Liang Tao · 2019
In this paper, we are interested in the conversion of whispered to normal speech. The baseline method uses standard bidirectional LSTM (BLSTM) RNN to predict both the spectral features and excitation parameters of the normal speech from whispered speech. Also, it employs STRAIGHT speech synthesizer. The BLSTM based whispered speech to normal speech conversion system is among the best systems in term of the naturalness of generated speech. However, in many cases, the model complexity and inference cost of BLSTM prevents its usage. As opposed to using standard BLSTM with sharing values, we propose a meta-network to generate non-shared weights for LSTM memory block in BLSTM (denote as meta-BLSTM). Besides, we use a low-rank approximation to generate the parameter matrix, which can reduce the model complexity. To our knowledge, this is the first study that uses meta-network to train a whispered to normal speech conversion system. To evaluate the performance of the proposed system, we performed experiments in the TIMIT dataset. Experimental results show that the proposed method achieves state-of-the-art performance.