Delve Into Neural Activations: Toward Understanding Dying Neurons

Ziping Jiang, Yunpeng Wang, Chang‐Tsun Li, Plamen Parvanov Angelov, Richard Jiang · IEEE Transactions on Artificial Intelligence · 2022

Theoretically, a deep neuron network with nonlinear activation is able to approximate any function, while empirically the performance of the model with different activations varies widely. In this work, we investigate the expressivity of the network from an activation perspective. In particular, we introduce ageneralized activation region/patternto describe the functional relationship of the model with an arbitrary activation function and illustrate its fundamental properties. We then propose a metric namedpattern similarityto evaluate the practical expressivity of neuron networks regarding datasets based on the neuron level reaction toward the input. We find an undocumenteddying neuronissue that the postactivation value of most neurons remain in the same region for data with different labels, implying that the expressivity of the network with certain activations is greatly constrained. For instance, around 80% of postactivation values of a well-trained Sigmoid net or Tanh net are clustered in the same region given any test sample. This means most of the neurons fail to provide any useful information in distinguishing the data with different labels, suggesting that the practical expressivity of those networks is far below the theoretical. By evaluating our metrics and the test accuracy of the model, we show that the seriousness of thedying neuronissue is highly related to the model performance. At last, we also discussed the cause of thedying neuronissue, providing an explanation of the model performance gap caused by the choice of activation.

Read the paper · More papers on PaperTik