The relative importance of static versus spectral change acoustic features for automatic speaker identification

Stephen A. Zahorian, Peter Guzewich, Xiao Chen, Roozbeh Sadeghian, Hao Zhang · The Journal of the Acoustical Society of America · 2017

For at least two decades, the primary acoustic features used for both automatic speech recognition (ASR) and automatic speaker identification (SID) have been Mel frequency cepstral coefficients (MFCCs) and their first and second order difference terms, referred to as Delta and double Delta terms. The MFCC’s capture static spectral information, whereas the Delta terms capture spectral change information. In this experimental paper, we first reformulate the MFCC’s and Delta terms, as discrete cosine transform coefficients (DCTCs), which take the place of the MFCC’s, and discrete cosine series coefficients (DCSCs), which take the place of the Delta terms. Low dimensionality DCSC spaces (spectral change) result in very poor speaker discriminability, as compared to discriminability based on DCTCs. However, reasonably accurate automatic speaker identification can be achieved in a high dimensionality DCSC space. Combining DCTC terms and DCSC terms results in only modest improvements in identification accuracy over what can be achieved with DCTC terms alone. We conclude, for the purposes of automatic speaker identification, static spectral information is far more informative than spectral change information. The results of this study, plus results in the literature, support the hypothesis that a similar conclusion could be reached for human ability to recognize speakers.

Read the paper · More papers on PaperTik