Singers Voice Identification and Authentication based on GFCC using K-means Clustering and DTW
Iosr Journals, Swati Atame, Shanthi Therese S. · Figshare · 2015
The use of computers has increased to a very large extent due to its immense use and growing technologies that gives rise to transfer of digital media over longer distances. As the audio data may travel over large networks it needs to be purified or cleaned during recording or while the processing the data so that it can be identified. Cleaning is related to removing the noise that may exist in the signal. There are 2 ways in which the audio can be cleaned. First, recording can be done in clean environment that includes no noise at all, even the fans may create disturbances, recording done in a closed room. But feature extraction in this case is more suitable using MFCC. Second, if recording is already done and needs to processed and if it contains lot of noise than some noise reduction technique such as wavelet or filter can used to eliminate the noise. In this paper, we will implement Gamma tone Frequency Cepstral Coefficient (GFCC) to eliminate the noise components. This is due to the fact that MFCC does not give required accuracy in the presence of noise. It works well in clean environment. On the other hand, GFCC gives the required performance and accuracy in clean as well as in the presence of noise. This paper compares auditory feature based methods namely GFCC and MFCC and concludes which method is more suitable in which environment. This resulting GFCC features are used by k- means clustering to group the similar voices in to clusters and thus identify the singer. Dynamic time warping is used to align or match two sequences that may vary in time or speed. It is a pattern matching technique that will authenticate the person a claimed identity.