A Romanian corpus for speech perception and automatic speech recognition

Ahsanul Kabir, Mircea Giurgiu · International Conference on Electronics, Hardware, Wireless and Optical Communications · 2011

A speech corpus is available in Romanian to use as the common material in speech perception and automatic speech recognition. It consists of high-quality audio of 400 sentences spoken by each of 12 speakers. Utterances are simple, syntactically identical phrases such as muta bronz cu p 2 agale. Preliminary intelligibility tests using the audio signals suggest that the collected speech is easily identifiable in quiet and low levels of noise. The corpus is annotated at the phoneme, syllable and word level and is available on the website for research use.

Read the paper · More papers on PaperTik