A Brazilian Speech Database
Marco Aurelio Deoldoto Paulino, Yandre M. G. Costa, Alceu S. Britto, Alisson Renan Svaigen, Linnyer Beatrys Ruiz Aylon, Luiz S. Oliveira · 2018
This work introduces a Brazilian Speech Database (BrSD), a novel dataset freely available created to support the development of speech-based recognition tasks. As far as we know, this is the first Portuguese language based database with these characteristics created and made available to the research community. We also describe experiments accomplished on BrSD exploring its different possibilities of classification tasks, i.e., age group and gender classification. We use four well-known acoustic features extracted directly from the audio signal and one texture-based feature extracted from a visual representation of the audio signal, the spectrogram. We considered three different classification scenarios: each feature individually, early fusion of the features, and late fusion of the features. Experiments were conducted using Support Vector Machine (SVM) and Multi-layer Perceptron (MLP) classifiers. The obtained results showed that SVM classifier achieved the best recognition rates both in early and late fusion scenarios. The best recognition rates achieved were 91.25%, 88.75%, and 80.25% for gender, age group, and age-gender classification tasks, respectively.