WoC SVC: A Model for Enhancing Lullabies Through Personalized Voice Cloning

Sanika Ardekar, Riddhi Sanghani, Sahil Nair, Ruhina B. Karani · Procedia Computer Science · 2025

This paper introduces the Whisper Of Comfort Singing Voice Conversion (WoC SVC) model, a system designed to transform lullabies sung by any voice into a similar, familiar voice, harnessing the positive impact of familiar voices on infants’ development and well-being. Central to the WoC SVC model is the integration of UnivNet, an advanced neural vocoder architecture, which facilitates the generation of high-quality audio waveforms while preserving the original content and melody. The system architecture comprises essential components, including a Content Encoder for extracting mel-spectrograms and log loudness features, a Soft Content Encoder (ContentVec) for generating compact content representations, and a Speaker Embedder for capturing target voice characteristics. These representations are concatenated and processed through UnivNet to synthesize audio waveforms in a similar, familiar voice. Comparison between SOVITS SVC model and WoC SVC is made and the result is evaluated using the values from MFCC cosine similarity and PESQ scores where WoC SVC has consistently performed better. The PESQ score for WoC SVC is 1.4929 and for SOVITS SVC is 1.1348, indicating a percentage difference of approximately 31.53%. Similarly, the cosine similarity for WoC SVC is 0.9346124 and for SOVITS SVC is 0.8947109, with a percentage difference of around 4.46%. These results demonstrate that WoC SVC outperforms SOVITS SVC.

Read the paper · More papers on PaperTik