Stochastic Filter Techniques in Combination with Sliding Mode Differentiators as the Basis for a Reliable Neural Network-Based Recognition of Phonemes in Speech Signals

Andreas Rauh, Matthew Schmidt, Susann Tiede, Cornelia Klenke · 2018

The automatic estimation of the fundamental frequencies of phonemes is one of the important building blocks for the recognition of pronunciation disorders in spoken language. Based on the estimation of the so-called formant frequencies, it becomes possible to distinguish between voiced and unvoiced phonemes and to identify them reliably. Both the filter-based frequency estimation and the phoneme classification by neural networks are tasks that are investigated in the research project SUSE (A Software assistance system for Uncovering speech disorders by Stochastic Estimation techniques) bringing together the fields of signal processing as well as speech therapy and phonology. In this paper, a stochastic frequency estimation scheme based on the Unscented Kalman Filter is, firstly, extended by a sliding mode differentiator to enhance the accuracy of frequency estimation in naturally spoken language. Secondly, the estimation results are employed to train and implement a fundamental neural network classifier that can be used to distinguish automatically between different voiced and unvoiced phonemes. Classification results for an excerpt from a TV news broadcast conclude this paper.

Read the paper · More papers on PaperTik