Rekhta: An Open-Source High-Quality Urdu Text-to-Speech Synthesis Dataset

Abdul Rehman, Mahnoor Mehmood, Hammad Bakhtiar, Moiz Ahmed, Muhammad Umair Arshad, Naveed Ahmad · 2023

A dataset is one of the most pivotal components in creating and developing Deep Learning and Machine Learning models. It is fundamental to the idea of training, refining, and evaluating AI models, serving as a metric for evaluating the model's results. As an outcome of advancements in innovation, research, and AI analysis, datasets in numerous languages and formats have emerged on the internet. However, the expanse of internet resources appears to contract when confronted with Machine Learning and Deep Learning tasks in Urdu. The sheer availability of resources and authenticated datasets in the Urdu language, juxtaposed with a multitude of subpar, incomplete, and unverified datasets, significantly hinders the advancement of AI in the Urdu language. This paper documents the exploration and refinement of the Common Voice Urdu Corpus dataset version 12.0 to create a clean and refined dataset suitable for training Urdu Text-to-Speech models. This task has been accomplished through primary research, preliminary analysis, techniques such as fuzzy string matching through Levenshtein Distance, and meticulous comparison among various models.

Read the paper · More papers on PaperTik