Recuperative Reference Search Aided Approach Towards Voice Cloning

Premanand Pralhad Ghadekar, Ojas Joshi, Jignesh Barhate, Basun Kundu, Arman Nandeshwar, Manthan Gujar · 2024

This proposed system presents a novel approach to enhancing synthesized speech by infusing emotional nuances. By fine-tuning existing models, emotional inflections and intonations are introduced, improving perceptual evaluation of speech quality (MFCC) and mean opinion scores (MOS). The paper presents the solution of cloning life-like voice of any individual to an astounding similarity to that of original which is achieved in vocal data as small as 10 minutes in duration. The paper sees the contributions made such as use of deep invariant optimal quantization to effectively condense neural network models without compromising their performance, surpassing results in the field of RPA, RCA and OA by factors of $\mathbf{6. 2 1 \%}$, 6.9% and 2.12% respectfully with the added. The collaborative effort across machine learning, natural language processing (NLP), and audio engineering promises high-quality audio indistinguishable from human speech, with applications spanning assistive technologies to customer service solutions. This research contributes to advancing voice cloning technology, ensuring synthesized speech integrates emotional depth and authenticity effectively.

Read the paper · More papers on PaperTik