Decoupling Structural and Quantitative Knowledge in ReLU-based Deep Neural Networks

José F. Duato, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana–Ort́ı, José Cano · 2025

The relentless growth of artificial intelligence applications has led to substantial economic and environmental costs associated with training deep neural networks (DNNs). Recognizing the challenges in further optimizing conventional DNN training, in this paper we propose a novel approach that decouples structural information (non-linear functions) from quantitative knowledge (model parameters), and provide strong experimental evidence to demonstrate that these two types of knowledge can be trained independently. We evidence that ReLU-based DNNs can be deployed as globally linear models, from which different parts of the DNN are active for each sample, thus emulating the piece-wise linear outputs generated by ReLU activation functions. Leveraging this linear model foundation, this kind of DNN supports various objectives, including faster re-training times and combining multiple copies trained on different datasets for incremental and federated re-training.

Read the paper · More papers on PaperTik