Personalized Audiobook Generation Using Voice Cloning and Text-to-Speech Synthesis

Chun-Cheng Wei, Long-Chen Hung, Siyu Gao, Yi-Chang Chen, Shih-Pang Tseng · 2024

This project leverages advancements in voice cloning and text-to-speech (TTS) to develop a novel audio converter for personalized audiobook creation. Users upload a brief voice sample, allowing the system to learn and replicate unique vocal characteristics, such as tone and pitch, producing customized audio outputs. Integrating voice cloning with natural language processing (NLP) and TTS, the system first analyzes text to identify characters, dialogues, and emotional cues. This enables realistic, expressive narration tailored to each character and context. Unlike traditional TTS, this approach reduces monotony, allowing users to "narrate" stories in their own voice, creating an immersive and individualized listening experience.

Read the paper · More papers on PaperTik