Application of Speech Feature Extraction Method Based on Deep Learning in Speech Cloning

Jing Zhang, Yiyao Chen, Hongbin Ma · 2024

Although voice cloning technology has been widely used, there are still problems such as insufficient accuracy, lack of humanization, and even difficulty in reproducing the emotions of the original text. Therefore, this article believed that it is necessary to extract speech features in order to enrich the cloned objects of speech cloning and pursue higher speech restoration. Moreover, this article also intended to assist in this process based on deep learning technology. This article also conducted a comparative test at the end, comparing the performance of the conventional voice cloning process and the process based on deep learning technology in terms of error rejection rate and error acceptance rate. The final result was that these two indicators for conventional processes were 13.59% and 16.34%, respectively, while for processes based on deep learning technology, they were 9.99% and 13.21%, respectively. It can be seen that after using deep learning, the voice cloning process has significantly reduced the error rate. Therefore, this study found that from the perspectives of prosodic imitation and speech features, speech cloning based on deep learning technology can better restore in these two aspects, significantly making the cloned speech closer to the original speech.

Read the paper · More papers on PaperTik