Time-frequency-energy representation based real-time speech recognition

Danke Wu, J.N. Gowdy · 2002

The authors present an approach to isolated-word speech recognition which is characterized by two aspects: (1) nonlinear time normalization based on the gradients of short-time energy in a specific number of frequency bands, which retains the transient portions and ignores the steady-state portions of the speech signal in the frequency domain; and (2) real-time implementation due to low computational load. Simulation has shown that the correct rate of recognition was 99.5% for multiple speakers based on the TI-20 speech database. A very high accuracy for on-line recognition was also obtained.>

Read the paper · More papers on PaperTik