Text-to-Cadence: Synthesizing Rhythmic Voice Through Tacotron2 and Waveglow
P Kumar, Senthil Pandi S, Muqaddam Aaqil Sheriff, Sathish Kumar Kannaiah · 2024
The voice represents a complex mechanism employed by humans for the purpose of communication. As human beings, we possess the remarkable capacity to articulate our thoughts through vocalization, or more accurately, to manipulate molecules in a manner that conveys meaning. Each individual possesses a distinct quality to their voice, an attribute that remains beyond our control. Conventional approaches have predominantly concentrated on the advancement of text-to-speech models, attaining exemplary performance levels. This article introduces an innovative methodology, text-to-cadence, designed to elevate the level of personalization. We are pleased to present text-to-cadence, a sophisticated method that takes both text and the desired artist's name as input, generating an audio file that features their rhythmic vocal interpretation of your text as output.