Speaker recognition via fusion of subglottal features and MFCCs
Harish Arsikere, Hitesh Anand Gupta, Abeer A. Alwan · 2014
Motivated by the speaker-specificity and stationarity of subglot-tal acoustics, this paper investigates the utility of subglottal cep-stral coefficients (SGCCs) for speaker identification (SID) and verification (SV). SGCCs can be computed using accelerom-eter recordings of subglottal acoustics, but such an approach is infeasible in real-world scenarios. To estimate SGCCs from speech signals, we adopt the Bayesian minimum mean squared error (MMSE) estimator proposed in the speech-to-articulatory inversion literature. The joint distribution of SGCCs and speech MFCCs is modeled using the WashU-UCLA corpus (containing simultaneous recordings of speech and subglottal acoustics), and the resulting model is used to obtain an MMSE estimate of SGCCs from unseen (test) MFCCs. Cross-validation experi-ments on the WashU-UCLA corpus show that the estimation ef-ficacy, on average, is speaker dependent. A score-level fusion of MFCC and SGCC systems outperforms the MFCC-only base-line in both SID and SV tasks. On the TIMIT database (SID), the relative reduction in identification error is 16, 40 and 51% for G.712-filtered (300–3400 Hz), narrowband (0–4000 Hz) and wideband (0–8000 Hz) speech, respectively. On the NIST 2008 database (SV), the relative reduction in equal error rate is 4 and 11 % for 10 and 5 second utterances, respectively. Index Terms: speaker recognition, subglottal acoustics, cep-stral coefficients, score combination, MMSE estimation