Identifying Sub-documents in a Composite Scanned Document Using Naive Bayes, Levenshtein Distance and Domain Driven Knowledge Base

Oluwatosin Ogundare, Nathanial Wiggins · 2018

This paper presents an approach to classification of scanned documents by identifying the title of the document using Naïve Bayes classifier, the predicted title is then validated using a domain driven knowledge base. Levenshtein distance is used to mitigate errors arising from the OCR (optical character recognition) algorithm. This approach produced significantly better results than using the Naïve Bayes classifier by itself. This study contributes resources to the intelligent processing of real estate documents in the form of rich domain specific knowledge base.

Read the paper · More papers on PaperTik