Deep Neural Networks with ReLU-Sine-Exponential Activations Break Curse of Dimensionality in Approximation on Hölder Class
Yuling Jiao, Yanming Lai, Xiliang Lu, Fengru Wang, Jerry Zhijian Yang, Yuanyuan Yang · SIAM Journal on Mathematical Analysis · 2023
Abstract. In this paper, we construct neural networks with ReLU, sine, and [Formula: see text] as activation functions. For a general continuous [Formula: see text] defined on [Formula: see text] with continuity modulus [Formula: see text], we construct [Formula: see text]-sine-[Formula: see text] networks that enjoy an approximation rate [Formula: see text], where [Formula: see text] are the hyperparameters related to widths of the networks. As a consequence, we can construct [Formula: see text]-sine-[Formula: see text] network with the depth 6 and width [Formula: see text].[Formula: see text] that approximates [Formula: see text] within a given tolerance [Formula: see text] measured in the [Formula: see text] norm with [Formula: see text], where [Formula: see text] denotes the Hölder continuous function class defined on [Formula: see text] with order [Formula: see text] and constant [Formula: see text]. Therefore, the [Formula: see text]-sine-[Formula: see text] networks overcome the curse of dimensionality in an approximation on [Formula: see text]. In addition to its super expressive power, functions implemented by [Formula: see text]-sine-[Formula: see text] networks are (generalized) differentiable, enabling us to apply stochastic gradient descent to train.