Application of Three Different Artificial Neural Network Architectures for Voice Conversion

Bageshree Sathe-Pathak, Shalaka Patil, Ashish R. Panat · Advances in intelligent systems and computing · 2016

This paper designs a Multi-scale Spectral transformation technique for Voice Conversion. The proposed algorithm uses Spectral transformation technique designed using multi-resolution wavelet feature set and a Neural Network to generate a mapping function between source and target speech. Dynamic Frequency Warping technique is used for aligning source and target speech and Overlap-Add method is used for minimizing the distortions that occur in the reconstruction process. With the use of Neural Network, mapping of spectral parameters between source and target speech has been achieved more efficiently. In this paper, the mapping function is generated in three different ways, using three types of Neural Networks namely, Feed Forward Neural Network, Generalized Regression Neural Network and Radial Basis Neural Network. Results of all three Neural Networks are compared using execution time requirements and Subjective analysis. The main advantage of this approach is that it is speech as well as speaker independent algorithm.

Read the paper · More papers on PaperTik