A Roberta-based model for identifying non-substantive factual elements of the case
Sijia Huang, Yu Shan Chen, Enji Zhou, Zilong Zhuan · 2021
Legal documents contain rich information of case elements, and the automatic extraction of case elements can effectively promote the development of China's judicial intelligence field. In this paper, a pre-trained multi-label text classification model based on the RoBERTa model domain is designed to identify non-substantive factual elements in legal documents. The model uses weighted summation in the text encoding layer to get the RoBERTa encoding output and employs a multi-headed attention mechanism to capture factual descriptions and association information between elements. The critical fact elements are identified by Transformer secondary encoding and CNN extraction of critical features. The experimental results at CAIL2019 show that the model proposed in this paper has substantial performance improvement compared to other baseline models, with Macro-F1 and Micro-F1 values improving by 5.04% and 6.43%, respectively to the traditional Roberta model on the three datasets. It indicates that the model can capture richer labeling features and thus effectively identify the non-substantive core elements of legal documents.