Detection of spectral transition for speech perception based on time-frequency analysis
Qun Chao Zhao, Qianli Gao, Huisheng Chi · 2002
Current speech or speaker recognition systems rely largely on voiced parts of utterance, though a great amount of information for speech perception is contained in the nonstationary consonants and transition. How to model and characterize the dynamic spectral features describing the transition still remains a question. This paper investigates the modeling and detection of the spectral transition based on time-frequency analysis. Linear and nonlinear modeling of the transitions are proposed using linear and quadratic frequency modulation signals. Then two strategies of detection of the spectral transition are presented, i.e., the Radon-Wigner transform (RWT) and Radon-ambiguity transform (RAT). Both simulated and real speech data from the TIMIT database are used to test the detection procedure.