Structure-based voiced/usable speech detection using state space embedding

Uchechukwu O. Ofoegbu, Brett Y. Smolenski, R.E. Yantomo · 2005

The process of speech production in the human system is very complex, possesses nonlinearities, and can only be precisely modeled in terms of nonlinear dynamics. A non-linear speech classification approach is proposed, which classifies speech based on features extracted from Takens' method of delays, a technique used to reconstruct signals into a trajectory in multidimensional state space. In this research, two types of speech detection are presented, namely, voiced and usable speech (for speaker identification purposes). The proposed approach has been able to yield a probability of error of 12% in noisy environments for voiced speech detection, and 78% correct usable speech detection by comparing the structures of embedded voiced speech frames with embedded unvoiced speech frames, and embedded usable speech frames with unusable speech. Some applications of this speech detection technique include the enhancement of speaker identification and speech recognition systems.

Read the paper · More papers on PaperTik