Comparative study of voice conversion framework with line spectral frequency and Mel-Frequency Cepstral Coefficients as features using artficial neural networks
Amit Kumar Bhuyan, Jagannath H. Nirmal · 2015
This paper is intended to formulate the mapping function using Feed-forward Neural Networks on Line Spectral Frequency and Mel Frequency Cepstral Coefficient and to compare their outcomes to decipher the better solution to the spectral mapping impediment. The experimentation is confined to the augmentation of spectral and excitation (glottal) domains of speech. LSF and MFCC are used to represent the spectrum and as input predictor variables to the above mentioned neural networks. It contains the use of neural network in the voice conversion framework. The function of artificial neural network is to map the spectral characteristics of a source speaker to the target speaker so as to obtain an authenticated voice conversion model. The temporal alignment of the speech uttered by source and the target is attained using Dynamic Time Warping (Dynamic Programming). The excitation mapping is accomplished using Residual Selection method. The performances of these Voice Conversion systems are assessed using subjective and objective measures which assure the genuineness of the conversion system design.