Interpretable Text Classification in Legal Contract Documents using Tsetlin Machines

Rupsa Saha, Sander Jyhne · 2022

Legal text contains various challenges in automated processing, compounded by the lack of detailed resources available for them. However, the ability of process such texts automatically is highly sought after. In this paper we try to parse a set of contract documents and identify key legal terminologies present in them, with the help of four text processing methods from different backgrounds: Tsetlin Machines, BERT, CNNBiLSTM and FastText. We show that the TM based approach works at par with other popular methods, with the added benefit of making available important clause literals that can act as specific linguistic cues to legal terminology.

Read the paper · More papers on PaperTik