Spa-L Transformer: Sparse-self attention model of Long short-term memory positional encoding based on long text classification
Shengzhe Zhang, Jiayu Ye, Qingxiang Wang · 2023
The emergence of Transformer and its derivative models brings new opportunities to tasks of NLP (Natural Language Processing). Transformer is not only a separate model, but also the core of different text task systems. Therefore, Transformer has become an important component of many powerful models. However, Transformer is not without defects. Researchers are still puzzled by the huge amount of computation generated in the process of self-attention. Especially in long text data sets. We propose a new ProbSparse self-attention Transformer model based for text classification. In the following, we will call it SpaL Transformer. We query important attention factors through KL divergence and add Long Short Term Memory(LSTM) to positional encoding, and only focus on the main query. Then, we select the most important relevant attention based on the confidence score to focus the overall attention. At the same time, we propose LSTM positive encoding to obtain relative position information to optimize the model. In long text dataset IMDB, our model improves the accuracy of Transformer. And F1 score improved by 0.064.