Reconstruction of the Face Image from Speech Recording: a Neural Networks Approach

Aleksey K. Bragin, Sergey A. Ivanov · 2021

Human speech can contain information about the personality, gender, emotional state of the speaker. In this article, we will explore the possibility of reconstructing an image of a human face from a recording of his voice. To achieve this goal, we design an autoencoder neural network and train it using data from AVSpeech, a dataset that contains videos of people talking, collected from the popular video hosting YouTube. To train the neural network, we select short pieces of audio containing speech and save a frame with the face of the corresponding person. The audio data is converted into a spectrogram, and images with human faces are cropped and normalized from the saved frames. Based on the data obtained, we train an autoencoder network using a self-supervised learning approach.

Read the paper · More papers on PaperTik