Separability of Human Voices by Clustering Statistical Pitch Parameters
Amit Banerjee, Sakshi Pandey, Mohammad Ali Hussainy · 2018
Speech is embedded with multiple information which includes the message to be transferred, the speaker information and the prosodic information. The physical properties of sound waves, such as mel-frequency cepstral and linear predictive cepstral coefficients are used in the speaker recognition system. While the prosodic features, like pitch and intensity, plays a vital role in human speech perception. In this paper, we investigate the speaker separability problem in real-time by clustering the statistical pitch feature. The goal is to identify the number of speakers in a given voice sample. For this, we analyze the space distribution of the prosodic parameter (pitch) to form clusters using the statistical parameters: mean, standard deviation, skewness and kurtosis. We experimentally observe that the statistical parameter belonging to a voice sample lie in the range of similarity. We calculate the variability for multiple speaker clusters and found that the intra cluster distance is 13 units, whereas the inter cluster distance varies according to the number of users in the voice sample.