Speech analysis in automatic speaker recognition schemes in both forensic and commercial applications
Javier Ortega-García, Joaquín González-Rodríguez · The Journal of the Acoustical Society of America · 2001
Speaker identity is a complex information included in the speech signal and codified in several levels of knowledge. Regarding the phonetic segmental level, this information resides in the specific acoustic resonances of the vocal tract that can be observed in the spectral envelope. Considering the nonstationarity of speech signals, short-time speech analysis techniques are employed through time-windowing procedures. Homomorfic cepstral analysis is then usually performed, as deconvolution between glottal information and vocal tract information is efficiently accomplished in this domain, guiding to a parametrization process of each frame through different techniques, as Mel-frequency-derived (MFCC) or linear prediction-based (LPCC) cepstral analysis. After the parametrization stage, probabilistic characterization of speakers identity is attained, usually through the use of left-to-right or ergodic hidden Markov models (HMM). Two different operating modes are clearly found in speaker recognition, namely (i) speaker identification, task in which a speaker has to be selected from a given set of speakers models, and (ii) speaker verification, in which a binary decision (accepted/rejected) has to be taken regarding the input unknown speech. Automatic speaker recognition leads to development of both commercial and forensic applications, but the cost of the decision of the system clearly separates the technical specificity of them.