Legal Text Retrieval with Contrastive Representation Learning and Evolutionary Data Augmentation

Youhua Zhou, Xueming Yan, Han Pang Huang, Haowen Yan, Minghao Chen · 2024

Legal text retrieval holds significant importance in the audit field, posing a challenge as a semantic matching problem. Despite the success of text semantic matching methods, particularly with the advent of large language models, these approaches face challenges when applied to domain-specific tasks, like legal text retrieval. Specifically, issues arise due to the concentrated distribution of data within the specific domain and the insufficient number of training samples. To address these challenges, this paper introduces a text semantic matching model tailored for the task of legal text retrieval, leveraging contrastive learning and evolutionary algorithms. A contrastive learning-based embedding model, which learns semantic representations in a feature space, is used to minimize the distance between matched text pairs and maximize the distance between unmatched text pairs. Additionally, an evolutionary algorithm-based sample augmentation model is introduced to augment the sample set and enhance the representational capabilities of the samples. The efficacy of the proposed method is evaluated in the context of legal text retrieval in the auditing field, and the experimental results reveal promising outcomes, with the proposed method achieving a Hits@l accuracy of 53.09%, a 2.99% improvement over the best baseline model. The Hits@20 accuracy reaches 75.15%, representing a 2.69% enhancement compared to the state-of-the-art methods.

Read the paper · More papers on PaperTik