Named Entity Recognition in Long Documents: An End-to-end Case Study in the Legal Domain

Hossein Keshavarz, Zografoula Vagena, Pigi Kouki, Ilias Fountalis, Mehdi Mabrouki, Aziz Belaweid, Nikolaos Vasiloglou · 2022 IEEE International Conference on Big Data (Big Data) · 2022

Named entity recognition (NER) is a fundamental task for several important applications such as knowledge base construction and semantic search. So far, the focus has been on building machine learning models, which identify generic named entities (e.g., person, date). Such models can be used off-the-shelf without requiring ground truth labels for training. However, such models cannot generalize to specialized domains that have domain-specific named entities (e.g., the legal domain). In these cases, it is inevitable to generate ground truth data and experiment with a variety of models in order to achieve good performance. Motivated by a real use case from the financial sector, we discuss the approach and lessons learned when solving the NER problem in the legal domain. This task is particularly challenging because it requires extensive human expertise to produce high quality ground-truth labels. For solving the legal domain NER problem, we first crawl a large dataset of legal documents and then introduce a semi-automated process to generate high-quality labels for a set of eleven predefined named entities. We validate that the proposed approach achieves high quality labels that outperform popular out-of-the-box NER methods. On top of that, our method once followed, can generate ground truth labels for the pre-defined named entities for an unbounded number of documents. Next, we experiment with a set of models and training procedures and report their performance on the NER task. Our experimental evaluation confirms that most of the models can generalize very well, achieving F1-score between 86% and 98.9%. The dataset, the labels produced by human annotators and our semi-supervised approach, as well as our code are made available to the research community.

Read the paper · More papers on PaperTik