Persian large vocabulary name recognition system (FarsName)

Alireza Hajitabar, Hossein Sameti, Hossein Hadian, Arash Safari · 2017

There has been no isolated word recognition database for the Persian language so far. In this paper we introduce FarsName dataset which contains 20 thousands isolated-word Persian utterances spoken by 226 speakers from all regions of the country each saying an average of 88 Persian names. There is a total of 5235 unique names in this dataset. Various cell phone brands have been used to record this dataset. This indicates the high diversity of the utterances in this dataset. We have been able to achieve 10.34% WER on this set using Kaldi. This is a very good performance considering the recording environment have been normal and potentially noisy.

Read the paper · More papers on PaperTik