AudioFacialMatrix: Dataset for Voice and Face AI

Rahul Singh, Rita Singh · 2024

This paper introduces the “AudioFacialMatrix” dataset, which is an amalgam of audio and visual data to aid research on human voice and facial imaging through Artificial Intelligence (AI) techniques. The dataset comprises 10,000 pairs of facial images and corresponding voice clips, organized by the nationality of individuals from eight different countries. In addition to detailing the dataset, this paper describes the methodology used for data generation and expansion. Lastly, a clustering task and a liquid neural network implementation are elucidated with a concluding discussion on potential methods to leverage this dataset.

Read the paper · More papers on PaperTik