Peer Review #2 of "ResMem-Net: memory based deep CNN for image memorability estimation (v0.3)"
2021
Image memorability is a hard problem in Image Processing due to the problem's subjective nature.But with the introduction of Deep Learning, availability of huge data and GPUs, great strides have been made feasible in predicting the memorability of an image.This paper proposes Residual Memory Network (ResMem-Net), a novel deep learning architecture that is a hybrid of Long Short-Term Memory (LSTM) and Convolutional Neural Networks (CNN).This architecture uses information from the hidden layers of the CNN, which are the learned convolution filters that contain extract features from the input, to compute the memorability score of an image.The intermediate layers are important for predicting the output because they contain information pertaining to the intrinsic properties of the image.The proposed architecture automatically learns visual emotions and saliency, as shown by the heatmaps generated using the Gradient Regression Activation Map (GradRAM) technique.The study also used the heatmaps and results to analyze and answer one of the most important questions in image memorability which is "What makes an image memorable?".The proposed model is trained and evaluated using the publicly available Large-scale Image Memorability dataset (LaMem) from MIT.The result shows that the proposed model achieves a rank correlation of 0.679 and a mean squared error of 0.011, which is better than the current state-of-the-art models and is close to human consistency (p=0.68).The proposed architecture also has a significantly low number of parameters compared to the state-of-the-art architecture, making it memory efficient and suitable for production.