Integration of Rule-Based Reasoning and Transfer Learning in Legal Document Review

Robert Keeling, Ava Guo, Peter Gronvall, Nathaniel Huber-Fliflet, Jianping Zhang · 2022 IEEE International Conference on Big Data (Big Data) · 2022

Protecting privileged communications and data from disclosure is paramount for legal teams. Unrestricted legal advice, such as attorney-client communication or litigation strategy, is exempt from disclosure in litigation or regulatory events and is vital to the attorney-client relationship. To protect this information from disclosure, companies and outside counsel must review vast amounts of documents to determine those that contain privileged material. This process is costly and time consuming. As data volumes increase, legal counsel employs methods to reduce the number of documents requiring review while balancing the need to ensure the protection of privileged information. Keyword searching is a popular method to target privileged information and reduce document review populations. Keyword terms are effective at casting a wide net but generally return overly inclusive results – most of which do not contain privileged information. To overcome the weaknesses of keyword searching, legal teams have started using supervised learning techniques to more precisely target privileged information. However, reviewing and labeling training documents is costly and time intensive and may cause counsel to forego the use of supervised learning in certain scenarios. In addition, supervised learning techniques may not find all the privileged documents in a document review and require companies to use keyword terms to identify critical privileged information. In this paper, the authors propose a novel method to automatically identify privileged documents without the need to label new training documents. This method integrates rule-based reasoning with transfer learning. Experimental results show that the proposed integrated method performs better than rule-based reasoning and transfer learning individually and can effectively identify privileged documents.

Read the paper · More papers on PaperTik