New resources for Brazilian Portuguese: Results for grapheme-to-phoneme and phone classification

Chadia Nadim Aboul Hosn, Luiz Alberto Baptista, Tales Imbiriba, Aldebaro Barreto da Rocha Klautau Junior · 2006

Speech processing is a data-driven technology that relies on public corpora and associated resources. In contrast to languages such as English, there are few resources for Brazilian Portuguese (BP). Consequently, there are no publicly available scripts to design baseline BP systems. This work discusses some efforts towards decreasing this gap and presents results for two speech processing tasks for BP: phone classification and grapheme to phoneme (G2P) conversion. The former task used hidden Markov models to classify phones from the Spoltech and TIMIT corpora. The G2P module adopted machine learning methods such as decision trees and was tested on a new BP pronunciation dictionary and the following languages: British English, American English and French.

Read the paper · More papers on PaperTik