A study of automatic speech recognition in Portuguese by the Brazilian General Attorney of the Union

Rodrigo Fay Verqara, Paulo Henrique dos Santos, Guilherme Fay Verqara, Fabio Lucio Lopes Mendonca, Carlos Eduardo Lacerda Veiga, Bruno J. G. Praciano, Daniel Alves Da Silva, Rafael T. de Sousa · 2022 IEEE International Conference on Data Mining Workshops (ICDMW) · 2022

This article presents a study of an automatic speech recognition system in Portuguese applied to videos by the General Attorney of the Union of Brazil. As they are confidential videos, using proprietary software from large companies is not allowed for security reasons. Thus, constructing an artificial intelligence model capable of performing automatic speech recognition in Portuguese in the judicial context and making this model available for large-scale inference is critical to maintaining data security. For this purpose, a dataset in Brazilian Portuguese was used by a combination of 3 datasets already built. The system used TDNN Jasper and QuartzNet architectures for network training, obtaining promising preliminary results, having a word error rate (WER) of 56% without using a linguistic model.

Read the paper · More papers on PaperTik