ALCG: Chinese Four-clause Compound Sentences Relation Classification Based on LERT Combining CNN and GRU
Zixuan Zhang, Yuan Li · 2024
Relation classification of compound sentences is an important task in Chinese information processing. Due to the extensive clause span, the intricate hierarchical structure, and the limited distribution in the corpus, the relation classification task for four-clause compound sentences presents significant challenges. In this paper, we propose the model ALCG, which integrates the linguistically-motivated pre-trained language model LERT, the convolutional neural network CNN, and the gated recurrent unit GRU. This model exhibits the current optimal performance on the dataset based on the CCCS (Corpus of Chinese Compound Sentence). To address the problem of the scarcity of the four-clause compound sentences training data, the model in this paper introduces the AEDA (An Easier Data Augmentation) module, which significantly improves the performance of the model in the case of using different sizes of training sets. Furthermore, the effect of training on only 70% of the training set exceeds the effect of training on full training set. This paper also examines the impact of a pre-trained language model on this task, and confirms the superiority of LERT in recognizing the kind of Chinese compound sentences.