WANI – A Text-To-Speech model for Indian Languages and Beyond for Aerospace Applications

Amresh Kumar, Harshil Gupta, Dhipu TM, R Rajesh · 2024

Introducing WANI, a Text-to-Speech (TTS) model designed to generate highly realistic, multilingual and lifelike audio. WANI utilizes a Transformer based architecture, leverages encodec and AudioLM. It is trained with over 1000 hours of audio data and provides immersive audio experience, especially for Indian languages. In airborne surveillance, where quick choices and effective communication are essential, use of WANI for data interpretation, multilingual communication and achieving overall operational effectiveness can prove to be of vital importance. The WANI system has obtained an average EER (Equal Error Rate) similarity score of 0.879 in the range of [-1, 1] with multilingual support.

Read the paper · More papers on PaperTik