Machine learning-based text to speech conversion for native languages
Rajashri Chitti, Shilpa Paygude, Kishor Bapat · 2023
Marathi is one of India's most ancient languages. This research project (TTS) focuses at the growth of the Marathi Text-to-Speech system. Marathi text produced in either Devanagari or Ekalipi script is used as input for Marathi TTS. The Text to Speech system's purpose is to convert any given text into speech. In order to mimic human speakers, speech synthesis involves creating technology that can produce human-like speech from any text input. Two key elements of a text-to-speech system are text processing and speech production. It is crucial that the text processing component provides an adequate sequence of phonemic units in order to construct a speech synthesis system that produces natural-sounding speech. The Hidden Markov Model is used with the statistical parameters for speech synthesis. This research is conducted to improve the pronunciation accuracy. The primary goal of this research is to improve and fine-tune the TTS system's pronunciation accuracy for native Indian languages like Marathi and make it seem more natural than current TTS systems (like how humans speak). The WASRAW ("Write as Spoken, Read as Written") technique is what the TTS system uses to speak written text. The objective is to create a language-independent solution as opposed to current solutions, particularly with regard to pronunciation. Unique script i.e., Ekalipi script has been used along with Devanagari script for Marathi language. The proposed model has been tested for 35 Marathi speakers for Devanagari script and achieved MOS of 4.38 for speech intelligibility, MOS of 4.45 for pronunciation accuracy and MOS of 4.44 for naturalness. Also, for Ekalipi script it has been tested for around 10 users who are familiar with the script. The dictionary-based approach has been followed in order to speak Ekalipi words. It has achieved good results for simple words written in Ekalipi. Even though the system is able to synthesize Ekalipi words correctly, there is still a need for a generalized model. As a future work, rules can be defined and the generalized model can be implemented for Ekalipi script and can be tested with the more no users and more no of languages can be added.