Show and Listen: Generating Music Based on Images using Machine Learning

Nirmal Mendis, R.M.N.S. Bandara · 2022

Music and images can be interpreted as arts of communicating emotions and feelings. Thereby, one can observe that there are instances where music and images are in a close association, such as an instance where background music is used to enhance the emotion depicted in a picture or video. A human composer would be able to compose such music by analyzing an image with the intention of sparking emotion in the subject within the context of the image. Thereby, the goal of this research is to develop a machine learning model which can perform the same task of composing a novel meaningful melody given an image. Inspired by image captioning using machine learning, the proposed architecture for the model involves using a Convolutional neural network (CNN) to extract image features and a Long short-term memory (LSTM) model to generate melodies. Following the model training, a subjective and objective evaluation was conducted and the obtained results indicated that the model performed well in accomplishing its goal. Naturally, though the model does not achieve perfection, it takes us one step closer to opening a wide range of possibilities for generating music from images using machine learning.

Read the paper · More papers on PaperTik