Sentient Sound waves: Elevating Emotional Communication with AI-Generated Speech Technology

Rahul Roshan G, Rohit Roshan, M H Sohan, S M Sutharsan Raj, V R Badri Prasad · 2024

This research-based project is about a new way to put feelings into computer-generated speech. It has two stages: text emotion detection and emotional speech synthesis. In the former part, labeled text data is used to build models to spot the emotions. This model learns to change pitch, volume, and rhythm based on the given emotion. During the latter phase, the text is changed into neutral speech using text-to-speech (TTS) methods. The model is used for practice to find the emotion in the text, tagging the text with that emotion. Also, a deep-learning model known as a tacotron is employed to make an emotional-sounding speech. The model perfectly uses machine learning strategies to blend feelings into speech. This opens up possibilities for turning plain text into speech with feelings of varied emotions. The proposed system makes a novice and important contribution to natural language processing for emotion identification and deep learning for emotional speech synthesis.

Read the paper · More papers on PaperTik