Activation Functions for Deep Learning Based on Generalised Entropies

R. A. Rudamenko, А. М. Savchenko, K. M. Semenov · Moscow University Physics Bulletin · 2026

The Shannon (Boltzmann–Gibbs) entropy is the foundation of classical statistical mechanics and deep learning; however, it encounters difficulties in describing the dynamics of non-extensive systems. In this paper, we propose the application of generalised entropies to construct new fundamental blocks for deep neural network architectures. The proposed approach generalises the classical softmax layer by employing the parametric entropies of Rényi, Tsallis, and Sharma–Mittal. The parameters $$q$$ and $$r$$ control the shape of the distribution: as $$q\to 1$$ , the optimal distribution converges to softmax, whereas at $$q=2$$ it converges to sparsemax. In particular, we consider a variant corresponding to $$q$$ -entmax, in which adaptivity is achieved by varying the parameter $$q$$ while keeping $$r$$ fixed. The study includes the derivation of analytical expressions for the Jacobian with respect to the parameters $$q$$ and $$r$$ for optimisation via explicit differentiation methods. A comparative analysis is carried out against existing approaches—softmax, sparsemax, and entmax (for $$q\in\{1.25,1.5,1.75\}$$ ). The results demonstrate improved performance metrics relative to softmax, sparsemax, and $$q$$ -entmax on a classification task with correlated class labels, leading to the conclusion that the Sharma–Mittal-based method is advantageous for this problem.

Read the paper · More papers on PaperTik