Employing the concept of stacking ensemble learning to generate deep dream images using multiple CNN variants

Lafta R. Al-Khazraji, Ayad R. Abbas, Abeer Salim Jamil, Zahraa Saddi Kadhim, Wissam Alkhazraji, Sabah Abdulazeez Jebur, Bassam Noori Shaker, Mohammed Abdallazez Mohammed, Mohanad Ali Mohammed, Basim Mohammed Al-Araji, Abdulkareem Z. Mohmmed, Wasiq Khan, Bilal Khan, Abir Jaafar Hussain · Intelligent Systems with Applications · 2025

The proposed work highlights the following major contributions. • As per the authors’ knowledge, current work is the first in applying stacking ensemble with multiple convolutional network (CNN) for deep dream. • Multiple octaves were used for the implementation of the stacked deep dream model. • Fine-tuning of the pre-trained CNN variants was thoroughly performed in order to obtain superior performance compared to other hybrid models. Addiction and adverse effects resulting from schizophrenia are rapidly becoming a global issue, necessitating the development of advanced approaches that can provide support to psychiatrists and psychologists to understand and replicate the hallucinations and imagery experienced by patients. Such approaches can also be useful for promoting interest in human artwork, particularly surrealist images. Accordingly, in the present, a stacking ensemble Deep Dream model was developed that aids psychiatrists and psychologists in addressing the challenge of mimicking hallucinations. The dream-like images generated in the present study possess an aesthetic quality reminiscent of surrealist art. For model development, a series of five pre-trained Convolutional Neural Network (CNN) architectures—VGG-19, Inception v3, VGG-16, Inception-ResNet-V2, and Xception were stacked in an ensemble learning approach to create Deep Dream images whereby the upper hidden layers of the architectures were activated, and the models were trained via the Adam optimizer. Performance of the proposed model was evaluated across three octaves to amplify the maximum possible patterns and features of the base image. The resulting dream-like images contain shapes that reflect elements from the ImageNet dataset on which the above pre-trained models were trained. Each of the base images was manipulated to generate various dreamed images, each one with three octaves, which were finally combined to construct the final image with its loss. Final Deep Dream image showed a loss of 47.5821, while still retaining some features from the base image.

Read the paper · More papers on PaperTik