Voice Over Vision: A Sequence-to-Sequence Model by Text to Speech Technology
E. G. Satish, P Ramesh Naidu, Girish Madhava Mogera, Harshvardhan karthik, Abhinandan · 2024
Text-to-speech (TTS) turns written text into spoken words using artificial voices. This uses natural language processing (NLP) and speech synthesis to make audio from text input. TTS has many uses - for people with visual impairments, hands-free communication, voiceovers for multimedia content. The development of neural network-based models like WaveNet and Tacotron has made synthetic speech much more natural and expressive. This paper reviews the history of TTS, the key components and techniques of modern TTS and future directions including emotion and personalization in synthetic speech. We also explore the ethical considerations and societal impacts of TTS and why we need to be responsible with its development and deployment.