The Multi-phase ReLU Activation Function
Chaity Banerjee, Tathagata Mukherjee, Eduardo L. Pasiliao · 2020
Deep Neural Networks have become the tool of choice for Machine Learning practitioners today. They have been successfully applied for solving a large class of learning problems both in the industry and academia, with applications in fields such as Computer Vision, Natural Language Processing, Big data Analytics and Bioinformatics. One important aspect of designing a neural network is the choice of the activation function to be used at the neurons of the different layers. Activation functions are used for introducing non-linearity into the neural network model so that the network can progressively learn more effective feature representations over the different layers. Several different activation functions have been studied and used in the literature, however, Linear, Sigmoid, Tanh and ReLU are the most commonly used activation functions. They are often selected empirically during the network design phase, rather than through a proper data driven process. In this work we study the problem of generalizing the single output ReLU activation to a multi-output variant using the idea of multi-phase ReLU where the parameters involved in the activation are learned through a data driven paradigm. We report results of experiments with the MNIST dataset using a two layer network that clearly demonstrates the efficacy of the multi-phase ReLU over the standard single output ReLU.