Improved Neural Network Algorithm Combining Adaptive Gradient Clipping and Self-Attention Mechanism
Sunjie Huang, Jun Xing, Yunfei Li · 2024
With the widespread application of deep learning in complex tasks, traditional neural networks face challenges such as gradient explosion and gradient vanishing when processing long-sequence data. Moreover, neural networks still have limitations in capturing global information and modeling long-term dependencies. To address these issues, this paper proposes an improved neural network algorithm that combines adaptive gradient clipping (AGC) with a self-attention mechanism. AGC dynamically adjusts the gradient clipping threshold based on the weights of different layers in the neural network, aiming to address the gradient explosion problem. The self-attention mechanism enhances the model's global representation capability by capturing long-term dependencies within the input sequence. To further optimize the model's performance, the paper also introduces optimization strategies such as multi-head self-attention, layer normalization, and mixed precision training. Experimental results demonstrate that the proposed algorithm excels across multiple tasks, significantly enhancing model training stability, convergence speed, and prediction accuracy.