Separate MAP adaptation of GMM parameters for forensic voice comparison on limited data
Chee Cheun Huang, Julien Epps, Ewald Enzinger · 2013
Automatic forensic voice comparison (FVC) studies have often employed Gaussian Mixture Model - Universal Background Model (GMM-UBM) modeling based on mean-only maximum a posteriori (MAP) adaptation or full MAP adaptation with little consideration of other variants of MAP adapted configurations. Our study indicates that an FVC system improvement in log-likelihood ratio cost (Cllr) of up to 6.8% can be achieved via fusion of other MAP adapted configurations such as variance-only and weight-only adaptations. We also demonstrate a novel adaptation and fusion strategy named Separate MAP (SMAP) that yielded a substantial FVC performance improvement in Cllr, up to 24.2%, and which is more robust under limited data conditions compared with the conventional mean-only or full MAP adaptation. The fusion strategy involves fusing multiple MAP adapted GMM sub-configurations where in each of these GMM sub-configurations only a small subset of the total number of GMM parameters are MAP adapted separately based on the same speaker-specific training data.