An Efficient Approach for Text-to-Speech Conversion Using Personalised Custom Voices
Kavya Chauhan, Chinmai Pore, Derek D’Souza, Pallavi Utkarsh Nehete, Megha A. Dhotay, Mrunal Aware · 2024
Text-to-speech is a speech synthesis in an Al system that converts text input to natural language speech. It is a genre of aiding technology that reads digital texts aloud. This technology can be executed by adding features of personalized voice models, which enable you to train a custom voice, constructing a unique voice as an output. This innovation is practical by making communication more user-friendly and efficient. The customer’s voice is recorded and used as training data sets. This training data will qualify for internal quality checks. The received datasets are trained into a custom voice model and further deployed. The testing data is applied to check the accuracy of audio outputs, ensuring the best audio delivery to the customer. These personalized voices are featured to define one’s mental & emotional state as well. This progress uses multiple key modules to achieve human-level quality on benchmark datasets.