Voices of Latin America: A low-resource TTS system for multiple accents

Jefferson Quispe-Pinares · 2024

In this article, we present the implementation of a low-resource Text-to-Speech (TTS) system with accents from various regions in Latin America. We evaluate different sources of databases such as freely accessible corpora, synthetic data, crowdsourcing, and audio recordings made in a professional studio by ourselves. We adopt the Cross-Industry Standard Process for Data Mining (CRISPDM) methodology for the development lifecycle of different models. The performance results of the TTS are represented by metrics such as Word Error Rate (WER) and Comparative Mean Opinion Score (CMOS) with interesting results like CMOS: Spanish:-1.13, Colombian:0.34, Mexican:-0.23. WER: Spanish: 15.54, Colombian: 15.65, Mexican: 16.1. Additionally, metrics are proposed to evaluate its performance in a conversational agent.

Read the paper · More papers on PaperTik