Non-symmetric time resolution for spectral feature trajectories
Stephen A. Zahorian, Jiang Yong Wu, Montri Karnjanadecha · The Journal of the Acoustical Society of America · 2011
In a study presented at the fall 2010 meeting of the Acoustical Society of America (Zahorian etal., “Time/frequency resolution of acoustic features for automatic speech recognition”), we demonstrated that spectral/temporal evolution features which emphasize temporal aspects of acoustic features, with relatively low spectral resolution, are effective for phonetic recognition in continuous speech. These features are computed using discrete cosine transform coefficients for spectral information from 8 ms frames and discrete cosine series coefficients (DCSCs) for their temporal evolution, over overlapping intervals (blocks) longer than 200 ms. These features are presented as an alternative to mel-frequency cepstral coefficients, and their delta terms, for automatic speech recognition (ASR). In the present work, it is shown that these features are even more effective for ASR, using a non-symmetric time window which is tilted toward the beginning of each block when computing DCSCs. This non-symmetry can be implemented by combining two Gaussian windows with different standard deviations. This work also supports the hypothesis that the left context is somewhat more informative to phonetic identity than is the right context. Experimental results for automatic phone recognition are given for various conditions using the TIMIT and NTIMIT databases.