Direct matching between music and image for contextual relationship analysis

Smart Wattanapornmongkol, Maethaphan Phaisalrattananukul, Phatthapol Saksermkiat, Parinya Sanguansat, Phanomyong Kaewprachum · 2024

We wanted to understand the relationship between emotions and objects or events. Thus, we have created a dataset of music-themed images and studied language context models to convert image, text, and audio data into vectors. We then developed two deep neural network models, Multi-Layer Perceptron (MLP) and CLIP-Based, to predict the most suitable song for any given image. It was found that MLP achieved an accuracy of 0.2057 when given both the image and its caption, compared to 0.2710 for CLIP-Based. However, when given only the caption, the accuracy of the CLIP-based neural network decreased to 0.2676, yet the accuracy slightly increase when given only the image to 0.2717. This suggests that direct image data could be enough to relate to songs, but providing descriptions of the images may not help the neural network to identify the relationships more accurately.

Read the paper · More papers on PaperTik