Personalized Audiobook Generation Using Voice Cloning and Text-to-Speech Synthesis
Chun-Cheng Wei, Long-Chen Hung, Siyu Gao, Yi-Chang Chen, Shih-Pang Tseng · 2024
This project leverages advancements in voice cloning and text-to-speech (TTS) to develop a novel audio converter for personalized audiobook creation. Users upload a brief voice sample, allowing the system to learn and replicate unique vocal characteristics, such as tone and pitch, producing customized audio outputs. Integrating voice cloning with natural language processing (NLP) and TTS, the system first analyzes text to identify characters, dialogues, and emotional cues. This enables realistic, expressive narration tailored to each character and context. Unlike traditional TTS, this approach reduces monotony, allowing users to "narrate" stories in their own voice, creating an immersive and individualized listening experience.