Word Representations for Neural Network Based Myanmar Text-to-Speech System
Aye Mya Hlaing, Win Pa Pa · International journal of intelligent engineering and systems · 2020
The main objective of this paper is to improve the naturalness of Myanmar Text-to-Speech (TTS) system without using time-consuming and expensive annotation of training corpus.Recently, word embedding, which has the advantage of training directly from a large amount of raw text data, has been used as the additional input features with the conventional input features or as the replacement of that conventional input features on the acoustic modelling of TTS systems.In this paper, the effectiveness of applying word vectors as the additional input features are investigated for Myanmar speech synthesis on the three acoustic modelling techniques, Deep Neural Network (DNN), Long Short-Term Memory Recurrent Neural Network (LSTM-RNN), and a hybrid of DNN and LSTM-RNN (Hybrid-LSTM-RNN).For the purpose of achieving better TTS performance, we built our own word vectors for Myanmar language.We further explore the best modelling method and vector dimension of word embedding for Myanmar TTS systems.Both objective and subjective evaluations are done on DNN, LSTM-RNN and Hybrid-LSTM-RNN based Myanmar TTS systems with and without additional input features such as Part-of-Speech (POS) and word vectors.According to the subjective results, applying additional input features in DNN, LSTM-RNN, and Hybrid-LSTM-RNN based Myanmar TTS systems can improve the naturalness of the synthesized speeches though objective results cannot lead to the significant improvement in LSTM-RNN and Hybrid-LSTM-RNN based systems.To the best of our knowledge, this is the first attempt to apply word vector features in neural network based Myanmar TTS systems.