Leveraging Document Structure for Better Classification of Complex Legal Documents

Alex Ratner · 2014

Document classification is a machine learning application that has been as impactful as it has been successful in a myriad of domains and applications. However, when the documents being classified are large and highly-complex, and when the set of potential classes is large as well, these models could be improved by incorporating more information about the documents’ overall structure. Most approaches use bag-of-words type models that discard local structure and focus on types of words or n-grams used. In this paper, we examine several models and attempt to leverage both local (e.g. n-gram) and global (e.g. structure and organization) document features. We apply these approaches to a new dataset of legal documents.

Read the paper · More papers on PaperTik