Interpretable Text Classification in Legal Contract Documents using Tsetlin Machines
Rupsa Saha, Sander Jyhne · 2022
Legal text contains various challenges in automated processing, compounded by the lack of detailed resources available for them. However, the ability of process such texts automatically is highly sought after. In this paper we try to parse a set of contract documents and identify key legal terminologies present in them, with the help of four text processing methods from different backgrounds: Tsetlin Machines, BERT, CNNBiLSTM and FastText. We show that the TM based approach works at par with other popular methods, with the added benefit of making available important clause literals that can act as specific linguistic cues to legal terminology.