A hierarchical two-level analysis structure for use in speech coding and recognition
T. Fjällbrant, F Mekuria · International Conference on Acoustics, Speech, and Signal Processing · 2003
The authors describe an analysis network that has been successfully applied to redundancy extraction in speech coding and for classification purposes in speech recognition. The network is designed as a two-level hierarchical structure with one building block. This building block performs a DCT (discrete cosine transformation) followed by magnitude and phase-derivative time trajectory computation. A combined short- and long-term analysis is obtained as well as energy compaction to a very small number of characteristic parameters. The parameters reflect both formant and pitch structure of voiced speech signals. The analysis applies to any quasi-stationary signal and is not dependent on any speech production model.>