Training Deep Neural Networks with HSIC and Backpropagation

Roshan Birjais, Kevin I‐Kai Wang, Waleed Habib Abdulla · 2024

Deep Learning has made significant strides in recent years, particularly in supervised learning tasks, leading to the development of numerous architectures aimed at improving various aspects of model performance. Despite the effectiveness of backpropagation (BP) and stochastic gradient descent in training deep networks, these methods are often hindered by time-intensive computations, exploding and vanishing gradients, and significant memory overhead. Alternative training strategies that reduce reliance on global BP are increasingly being explored to address these limitations. This paper proposes a simple architecture that integrates Hilbert Schmidt Independence Criterion (HSIC) layers with linear layers, where the HSIC layers are trained locally, and the linear layers are optimized using global BP. This hybrid approach mitigates the drawbacks of BP while enhancing the model’s ability to learn complex features across multiple layers. Our proposed model is benchmarked against existing HSIC-only models across several datasets, including MNIST, CIFAR-10, and Fashion MNIST. Results demonstrate the superior performance of our model, in terms of accuracy achieved and memory used. Additionally, we demonstrate its robustness and effectiveness when handling noisy data. The code is available at https://github.com/AreeBeee/HSIC.git

Read the paper · More papers on PaperTik