Call Translator with Voice Cloning Using Transformers
Nawaz Fateen Khan, Nandan Hemanth, Navyae Goyal, Pranav KR, Pooja Rani Agarwal · 2024
Voice Cloning employs technology and algorithms to create an artificial or synthetic reproduction of a person’s voice. To understand and mimic the distinct vocal qualities of the target speaker, including tone, pitch, cadence, and pronunciation, a machine learning model must be trained on audio samples of the speaker. This project presents itself as a revolutionary approach to the enduring problem of successful cross-lingual interactions, the paper’s novel approach to speech-to-speech machine translation that incorporates voice cloning. This method allows us to synthesize speech in the target language while maintaining the original speaker’s vocal characteristics. The model algorithms such as Whisper AI(WER 5% for transcription, No-Language-Left-Behind-200(44% more accurate than other models) for translation, Tacotron for speech synthesis and multiple transformers for voice cloning will be incorporated. The Voice Cloning model operates effectively on unseen voices, transcending the limitations of voice cloning will be incorporated. The voice cloning model operates effectively on unseen voices, transcending the limitations of relying solely on pre-trained models. The model successfully clones users’ voices in a different language with an approximate latency of 10 seconds.