LaFresCat: A Catalan Multi-Accent Speech Dataset for Text-to-Speech
Alex Peiró-Lilja, Martí Llopart-Font, Carme Armentano-Oller, Jose Giraldo, Ignasi Esquerra, Mireia Farrús, Baybars Külebi · 2024
Current generative text-to-speech (TTS) models are very robust and capable of learning the phonetics of a language almost perfectly. To do so, it remains crucial that the speech data used to train such models covers all phonetic richness. This includes phenomena of different accents. In the case of Catalan, although having access to various public speech corpora, there is a lack of high-quality, open access data covering its variety of accents. To meet this need, we have produced LaFresCat, a studio quality open-source Catalan multi-accent dataset with a total of 3.5 hours that covers 4 of the most prominent accents: Balearic, Central, North-Western and Valencian. We provide a detailed description of how utterances and recordings were produced. To evaluate the efficacy of LaFresCat, we trained a diffusionbased TTS model. Despite the small size of the dataset, we show that it is possible to generate accent-specific speech with an acceptable quality, and even enhance it by taking advantage of other Catalan datasets.