Review on Image Captioning and Speech Synthesis Techniques

K.V. Sruthi, M.S. Meharban · 2020

Caption generation may be a challenging AI problem to create a text description for a given picture. It needs to know the content of the image i.e., features and to transform these features of the image into words by a language model. Recently, deep learning techniques have accomplished results for this problem. The human-generated speech and the computer-generated speech are different. The quality of generated speech is defined by Intelligibility and naturalness. The quality of the audio generated is the Intelligibility and Naturalness is the quality of the speech generated, to judge the quality of generated speech. Deep Learning models demonstrate remarkably effective at learning the inherent features of information.

Read the paper · More papers on PaperTik