Dysarthric speech corpus in Tamil for rehabilitation research

Mariya Celin T A, T. Nagarajan, Parthasarathy Vijayalakshmi · 2016

This paper describes a speech data collected from 22 dysarthric speakers (7-female & 15-male) of various age groups in an Indian language, namely, Tamil. Dysarthric speakers who reported a diagnosis of cerebral palsy are chosen for speech data collection. The text for recording is chosen such that the influence of the articulatory errors can be observed in the initial, medial and end of the word. Each dysarthric speaker has uttered 262 sentences and 103 isolated words using a head mounted microphone. This corpus includes clinical data for all the 22 dysarthric speakers. The entire corpus is marked in sentence and word level. Time-aligned phonetic transcriptions are available for all the 22 speakers. The phonetic transcriptions are mapped with the IPA for a global use. The speech corpus is designed to act as a resource for the development of an automatic speech recognition system for speakers with neuromotor disability, research on articulatory dysfunctions, and rehabilitation research. The corpus includes unimpaired speech data, collected from 10 speakers (5-male & 5-female) for the same text used to collect dysarthric speech data along with the time-aligned phonetic transcription, for comparison. The corpus is available, on request, via secure FTP.

Read the paper · More papers on PaperTik