AN EFFECTIVE SUB-BAND BASED APPROACH FOR ROBUST SPEAKER VERIFICATION
Perasiriyan Sivakumaran, A. M. Ariyaeeinia, J. Hewitt, JA MALCOLM · 2024
The concept of splitting the entire frequency domain into sub-bands and processing these independently in between every consecutive recombination stage to generate a global score has already been investigated for speech recognition [|][2].Some aspects of this technique have also been studied for the task of speaker recognition [3][4].The main motivation for the above approach is that it allows for selective dephasis of sub-bands that are affected by narrow band noise and it permits emphasis of the sub bands which are more speci c to the speaker.It also provides the possibility of relaxing the conventional time-synchrony assumption between the sub-bands [l ][5].Moreover, the approach allows a closer simulation of the human perception [6].The main issue addressed in this paper is the reduction of the effects of any existing mismatch between the band-limited segments of the test and reference utterances in a sub-band basedspeaker veri cation system This can be achieved by using a weighting scheme which ensures that the scores associated with corrupted band-limited segments are appropriately deemphasised.The weighting factors requbed for this purpose can be computed using segmental scores obtained for a set of background speaker models.The general idea behind this approach is that if due to certain time and frequency localised anomalies there is some degree of mismatch between a particular band-limited segment of the test utterance (produced by the true speaker) and the corresponding segment of the target model.then a similar level of mismatch should exist between the considered test segment and the corresponding segments of the background speaker models.It is believed that through an appropriate selection of background speaker models, die above weighting scheme may lead to the emphasis of the sub-bands that are more speci c to the target speaker.The idea is based on the view that the mean separation between the scores of the target and background speaker models for a particular sub band is a measure of the performance of that sub band for the given target speaker.The paper also includes a study of two other aspects of the sub-band approach which have not been investigated for speaker recognition previously, These are the relaxation of the conventional time synchrony assumption of different sub-bands and the dif culties associated with the sub-band cepstral features.This paper is organised in the following manner.The next section details the classi cation process used in this work and describes the adopted merging strategy.Section 3 gives a description of the utilised speech database, and the method used for the extraction of sub-band feature vectors.The experimental work and results are detailed in Section 4. and the overall conclusions are presented in Section 5.