Formant Text To Speech Synthesis Using Artificial Neural Networks

Gurinder Kaur, Parminder Singh · 2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP) · 2019

Speech is a collection of words formed by voice. Simulation of human speech can be done by assuming fundamental frequency as glottis (source of sound) and the formant frequencies as the exacting arrangement of the speech organs, that is, on the location of the jaw, the tongue and the oral cavity. In this way, the human speech production system can be viewed as a source-filter model of speech production. This research has focused on to generate formant frequencies of any Punjabi sound and create a speech from a Punjabi text. The aim of this research is to synthesize speech waveform corresponding to input Punjabi text using formant synthesis technique. Python is used to create the model of speech generations and synthesis. The features of phonemes are extracted from recorded wave files, and those extracted features are stored in a database and generated formant frequencies for Punjabi sounds. Total 722 phonemes are taken for this research and 104 recorded wave files. The performance of the model was tested on the various parameters which are Formants, Fundamental frequency, voicing amplitude, Duration, Pitch, and Intensity. All these parameters were compared for better results and results indicated that these parameters in combination give better results.

Read the paper · More papers on PaperTik