Regularizing Temporal Explanations in Dynamic Neural Networks
Dalius Navakauskas, Martynas Dumpis · Electronics · 2026
Using attribution-based priors to improve the temporal interpretability and robustness of dynamic neural networks provides a computationally efficient method that does not alter the model structure during inference. We explore explanation-guided training for timeseries classification through the introduction of attribution-sensitive loss terms that serve as regularizers for the evolution of input relevance over time. The main contributions are the Temporal Relevance Smoothness Index (TRSI) and a ratio-based loss that reduces irregular step-to-step changes in channel-aggregated absolute relevance. TRSI is compared against temporal total-variation penalties computed using Layer-wise Relevance Propagation Total Variation (LRP-TV) and Integrated Gradients Total Variation (IG-TV). Experiments on a controlled three-class subset of the Korean University Human Activity Recognition (KU-HAR) dataset using a finite impulse response neural network (FIRNN) show that TRSI yields the strongest smoothness improvement, reducing the total variation of the aggregated relevance signal from 0.768 to 0.447 (41.8%), compared with 0.667 (LRP-TV) and 0.677 (IG-TV). Robustness tests indicate a clear advantage for TRSI under impulsive and white Gaussian test-time noise.