Enhancing Interpretability of Deep Canonically Correlated Autoencoders Through Convolutional Neural Networks and Visualization
Nipun Joshi · 2025
In this study, we studied unsupervised multiview learning techniques focused on maximizing correlation, particularly Deep Canonically Correlated Autoencoders (DCCAE). The goal of this study was to improve the interpretability of DCCAE. Two approaches were proposed i.e. one involved using a Convolutional Neural Network (CNN)-based DCCAE, and the other focused on visualization techniques. The CNN-based DCCAE outperformed the original feedforward neural network version. This improvement stemmed from CNN's ability to leverage pixel-level dependencies in MNIST data, enabling more effective learning of shared representations across views. This approach resulted in enhanced accuracy and reduced overfitting compared to the original model. To better understand and interpret the learned representations, three visualization techniques were applied i.e. Saliency Map, SmoothGrad, and Gradient-weighted Class Activation Mapping (GradCAM). These methods provided insights into how the model learned over time. The findings revealed that Saliency Map and SmoothGrad were more suitable for simpler network structures, while GradCAM proved particularly effective for more complex convolutional architectures. These advancements not only improved the performance of DCCAE but also made its internal workings more transparent and interpretable.