Machine Singing Generation Through Deep Learning
Risira Daksith Jayasinghe · 2021
This paper explores designing a singing synthesizer, that can generate any song once the lyrics and the musical notes are given. After training with pre-existing isolated vocal tracks of a singer, it would imitate the trained singer's style and voice when generating the vocals. The system is based on neural networks, and is designed to train on readily available vocal tracks. Flexibility to train with any singer's voice is a main feature, and this paper highlights methods incorporated to the system that allows to use training data from readily available sources without requiring manual transcribing. The synthesizer is based on a vocoder and three independent neural network models are designed to generate the required vocoder features corresponding to the pitch and timbre. Additional techniques that can be used to enhance the quality of the synthesized vocals are also explained in this paper. The entire system is packaged with a graphical user interface emphasizing the user-friendliness of this system. Evaluations through metrics and surveys are also provided in this paper to prove quantitatively and qualitatively that the system can provide satisfactory results.