Evaluation of Random-Projection-Based Feature Combination on Dysarthric Speech Recognition

Toshiya Yoshioka, Tetsuya Takiguchi, Yasuo Ariki · American Journal of Signal Processing · 2013

We investigated the speech recognition of persons with art iculation d isorders resulting fro m cerebral palsy. The articulation of their first speech tends to become unstable due to strain on s peech-related muscles, and that causes degradation of speech recognition. In th is paper, we propose a feature ext raction method based on RP (Random Pro jection) for dysarthric speech recognition. Random projection has been suggested as a means of space mapping, where the original data are projected onto a space using a random matrix. It represents a computationally simple method that approximately preserves the Euclidean d istance of any two points through the projection. Moreover, as we are able to produce various random matrices, there may be some possibility of finding a random matrix that gives better speech recognition accuracy among these random matrices. To obtain an optimal result fro m many random matrices, a vote-based combination is introduced in this paper. ROVER co mb ination is applied to the recognition results obtained from the ASR (Auto matic Speech Recognition) systems created from each RP-based feature. Its effectiveness is confirmed by word recognition experiments.

Read the paper · More papers on PaperTik