Enhancing Deep Neural Network Convergence and Performance: A Hybrid Activation Function Approach by Combining ReLU and ELU Activation Function

Ritesh Maurya, Divyam Aggarwal, T. Gopalakrishnan, Nageshwar Nath Pandey · 2023

Activation functions play an important role in Deep Neural Networks. The activation function can learn nonlinearities present in the data; therefore, it can learn intricate patterns present in the data. Rectified Linear Unit (ReLU) is an activation function that helps in encountering the problem of vanishing gradient. However, it suffers from ‘dying ReLU’ problem for the negative values. Leaky ReLU can solve the problem of ‘dying ReLU’; though it still suffers from a vanishing gradient problem due to the small gradient at for negative values, which results in slow convergence. Therefore, in this work, a combination of ReLU and Exponential Linear Unit (ELU) has been proposed considering the smoother convergence of the ELU activation function for the values on the negative side. Evaluating the effectiveness of the developed hybrid activation function compared to previous ReLU activation function versions such as SeLU, Leaky ReLU, ELU, etc. using a toy multi-layer perceptron and convolution neural network (CNN) model on the FashinMNIST and MNIST datasets. The improvement in the performance of these toy models when used with the proposed hybrid activation function on given datasets suggests the effectiveness of the proposed hybrid activation function.

Read the paper · More papers on PaperTik