Automatic recognition of ageing speakers
Finnian Kelly · Arrow@dit (Dublin Institute of Technology) · 2014
1580 2 e.g.The Xbox One games console, to be launched in November 2013, incorporates automatic speaker recognition for personalisation of the user-experience: http://www.xbox.com/en-US/xbox-one/get-the-facts#6-6 1 Chapter 6: Forensic speaker recognition and ageingVocal ageing is of particular relevance in the forensic domain.In this Chapter, the consequences of vocal ageing for forensic speaker recognition are investigated.The effect of vocal ageing on speech 'evidence' is established with a forensic automatic speaker recognition (FASR) system evaluation of the TCDSA males.A subsequent in-depth FASR evaluation of five Irish-accented ageing males provides a deeper insight.As forensic speaker recognition includes approaches based on subjective listening, the human perception of ageing change is of interest.The outcome Chapter 8: ConclusionsThe final Chapter draws together the main conclusions of this Thesis, and suggests possibilities for future study. Contributions of this thesisThis Thesis explores the effect of ageing on speaker recognition, and presents two new solutions to compensate for its negative effect on recognition accuracy.The contributions in each Chapter can be summarised as follows: Chapter 3Introduced the Trinity College Dublin Speaker Ageing (TCDSA) database -a new longitudinal speech database compiled for this study.Presented the largest longitudinal analysis of the acoustic correlates of ageing to date.Demonstrated the significant degradation in the performance of a Gaussian Mixture Model-Universal Background Model (GMM-UBM) speaker verification system due to the effects of long-term ageing. Chapter 4Introduced a stacked classifier approach for ageing speaker verification that exploits an ageingdependent decision threshold.Demonstrated this proposal to significantly reduce long-term classification error.Proposed a model-based quality measure, Wnorm, for speaker verification, and demonstrated that its ability to predict the verification scores of genuine-speakers could reduce classification error.Demonstrated that combining quality and ageing information in the stacked classifier framework reduces long-term classification error to a greater extent than either factor alone. Chapter 5Proposed eigenageing compensation for speaker verification, which learns information about the changes in the models of ageing speakers to compensate for ageing variability.Demonstrated that eigenageing compensation significantly reduces long-term classification error, and compares favourably to the stacked classifier approach.Proposed a new method of automatic age estimation using an ageing model learned as part of the eigenageing compensation procedure. Chapter 6Demonstrated that vocal ageing significantly undermines forensic automatic speaker recognition (FASR) by progressively weakening strength-of-evidence estimates.Applied eigenageing compensation to FASR, and demonstrated both its suitability for the domain and its ability to reduce the negative impact of ageing.Presented the outcomes of a listener test, and showed that vocal ageing becomes progressively more detectable with age difference and is significantly more detectable in female voices. Chapter 7Compared the performance of i-vector and GMM-UBM speaker verification systems in the presence of ageing.Demonstrated that the performance of both systems degrades at the same rate as ageing progresses.Concluded that dealing with ageing variability demands specific compensation strategies in addition to standard inter-session compensation approaches.1.