ReLU Neural Networks and Their Training
LUO Ge, X. G. Wang, Weizun Zhao, Sichen Tao, Zheng Tang · Mathematics · 2025
Among various activation functions, the Rectified Linear Unit (ReLU) has become the most widely adopted due to its computational simplicity and effectiveness in mitigating the vanishing-gradient problem. In this work, we investigate the advantages of employing ReLU as the activation function and establish its theoretical significance. Our analysis demonstrates that ReLU-based neural networks possess the universal approximation property. In addition, we provide a theoretical explanation for the phenomenon of neuron death in ReLU-based neural networks. We further validate the effectiveness of this explanation through empirical experiments.